0b41a0f1e7f66e0bce33b879275a93ccb9d8f976 lrnassar Mon Sep 28 16:33:18 2026 -0700 Fiber-seq description pages: bold the six data type names in the Compendium Description, and drop three passages that explained our own rendering rather than the data. The sample table no longer opens by justifying itself with how many checkboxes 41 samples times six data types would need; the difference track now says each position is colored by the most stringent threshold it meets, instead of describing the order the four signals are painted in; and the peak paragraph no longer explains that dense mode has no per-item hover. "Container name" and "subtracks" become "collection name" and "FIRE peaks", per the rule against exposing internal container terms, and the API paragraph points at the Table Browser for the peaks rather than only saying they are unavailable. Per Lou's review. refs #36210 diff --git src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html index 9d222bea464..3a0b84c7f99 100644 --- src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html +++ src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html @@ -1,48 +1,49 @@
This track is part of the Fiber-seq collection. It holds the full Fiber-seq data for 41 samples: 14 cell lines and 27 lymphoblastoid lines, 20 of them from individuals sequenced by the Human Pangenome Reference Consortium and five from rare disease cases. Chromatin accessibility and CpG methylation are read from the same molecules in the same experiment, so both are kept in one table here and can be compared without worrying about differences in cell preparation or sequencing depth. Six kinds of data are available for each sample:
-Because 41 samples times six kinds of data is far too many tracks for a checkbox list, samples -are chosen from a searchable table on this page. Pick the kinds of data you want along the top, -then select samples in the table; the browser turns on that combination for every sample you -picked, and keeps each sample's tracks together in the display. Sample class has filter -checkboxes in the panel to the left, and every column can be searched with the box under its -heading and clicked to sort, so a sample can be found by name, cell type, sample class or +Samples are chosen from the searchable table on this page. Pick the kinds of data you want along +the top, then select samples in the table; the browser turns on that combination for every +sample you picked, and keeps each sample's tracks together in the display. Sample class has +filter checkboxes in the panel to the left, and every column can be searched with the box under +its heading and clicked to sort, so a sample can be found by name, cell type, sample class or accession.
Accessibility, methylation and the haplotype overlays are all drawn 0 to 100 percent on a fixed scale, so heights are comparable between samples and between the two assays. The accessibility tracks use maximum as the windowing function, so a narrow element survives zooming out, while the methylation tracks use mean, since an average is the meaningful summary for a methylation level. FIRE peaks are shown in dense mode by default, one row per sample.
In both haplotype overlays: @@ -50,33 +51,32 @@
| Haplotype 1 | |
| Haplotype 2 |
The two are overlaid transparently, so a position with equal signal on both chromosomes appears as the two colors superimposed, and a haplotype-selective element appears as one color standing alone. Which parental chromosome is haplotype 1 is arbitrary and is not consistent between samples.
The CpG haplotype difference track runs from -100 to +100 percent, so a bar above the midline -means haplotype 1 is more methylated and a bar below it means haplotype 2 is. It is a stack of -four overlaid signals, one per significance threshold, drawn least significant first so that the -more significant levels are painted on top: +means haplotype 1 is more methylated and a bar below it means haplotype 2 is. Each position is +colored by the most stringent significance threshold it meets:
| All measured differences, regardless of significance | |
| p < 0.01 | |
| p < 0.001 | |
| p < 0.0001 |
The thresholds are nested, so a position drawn red also belongs to all three looser sets. Reading the track amounts to reading the color: grey is noise, red is a strong difference between the two chromosomes at that CpG.
@@ -96,33 +96,32 @@ GM25455, GM25456, GM27730, GM28570 and GM28572The classification is the one supplied by the laboratory, not one inferred from the cell type.
Peaks carry two filterable values, the FIRE score in the signalValue field and the false discovery rate as a -log10 value in the qValue field, and both can be filtered from a peak track's own configuration page, along with the score. No filter is applied by default. The FDR value stops at 100, which is the highest the pipeline reports rather than a real ceiling on significance, and 9 percent of the peaks in this track sit at it; filtering at the top of that range therefore selects a large group rather than a handful of outstanding peaks. A short tick inside each peak marks the point source, the single base the pipeline picked as the summit. -Switching a peak track to pack or full also gives each peak a mouseover with its FIRE score and -FDR; dense mode has no per-peak hover, which is a property of dense display rather than of this -track. The pValue field of the source files is set to -1 throughout and carries no information. +Switching a peak track to pack or full gives each peak a mouseover with its FIRE score and FDR. +The pValue field of the source files is set to -1 throughout and carries no information.
Permeabilized cells were treated with the Hia5 N6-adenine methyltransferase, which methylates adenines in DNA not protected by a bound protein, and high molecular weight DNA was prepared into PacBio SMRTbell libraries and sequenced. The adenine methylation added by the enzyme is chemically distinct from native CpG methylation, so both are read from the same molecule. Adenine methylation was called with fibertools-rs v0.4, and reads were aligned and haplotype-phased; for GM12878 an average 20 kb read spans at least one heterozygous variant and 87.9 percent of reads could be phased against GRCh38.
@@ -173,33 +172,33 @@ covering a single base, so its haplotype overlay drew nothing. The lab re-ran the sample with parental sequence added for phasing, and those tracks now carry full data. Its peak calls and CpG haplotype values changed with the reprocessing as well, so figures made from the first version of this track will not reproduce exactly for that sample.
The data can be explored interactively in table format with the Table Browser or the Data Integrator and exported from there to spreadsheet or tab-sep tables. From scripts, the data can be accessed through our API, one sample and data type at a time, for example track=fiberSeqCompendium_PM00001_acc for GM12878 accessibility or -fiberSeqCompendium_PM00001_cpg for its CpG methylation. The container name -fiberSeqCompendium holds no data of its own and cannot be queried, and the peak -subtracks are not yet available through the API. +fiberSeqCompendium_PM00001_cpg for its CpG methylation. The collection name +fiberSeqCompendium holds no data of its own and cannot be queried. The FIRE peaks are +not yet available through the API; use the Table Browser or the downloadable files for those.
For automated download and analysis, the data are stored in bigWig and bigBed files that can be downloaded from our download server, one directory per sample accession. Each directory holds the peaks as fire-peaks.ucsc.bb, the accessibility signal as all.percent.accessible.bw with hap1.percent.accessible.bw and hap2.percent.accessible.bw, and the methylation as cpg.combined.bw, cpg.hap1.bw, cpg.hap2.bw and four cpg.diffs_*.bw files. Individual regions or the whole genome annotation can be obtained using our tools bigBedToBed and bigWigToBedGraph, which can be compiled from the source code or downloaded as precompiled binaries for your system. Instructions for downloading source code and binaries can be found here. The tools can also be used to obtain features within a given range, e.g.