10f0f6d5160a96867ecf534e1b9d138df7d9b796
lrnassar
  Mon Sep 28 16:17:46 2026 -0700
Fiber-seq: hide the container by default, and split the five GM lines out
into a Rare disease sample class.  Max asked for superTrack on rather than
on show, since the track covers much the same ground as ENCODE DNase and
does not earn a slot in everyone's default hg38 view.  The five
lymphoblastoid lines GM25455, GM25456, GM27730, GM28570 and GM28572 had
been filed as Common Cell Line; Andrew Stergachis says they are rare
disease cases consented to broad genomic data sharing and the first of a
batch the lab intends to keep adding, so SAMPLE_CLASS_COLORS gains a third
entry and the facet now reads 20 HPRC, 16 Common Cell Line, 5 Rare disease
sample.  refs #36210

diff --git src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html
index 8ce5d3b224d..9d222bea464 100644
--- src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html
+++ src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html
@@ -1,245 +1,250 @@
 <h2>Description</h2>
 
 <p>
 This track is part of the <a href="hgTrackUi?db=hg38&amp;g=fiberSeq">Fiber-seq</a> collection. It holds the
 full Fiber-seq data for 41 samples: 14 cell lines and 27 lymphoblastoid lines, 20 of them from
-individuals sequenced by the Human Pangenome Reference Consortium and the rest from ENCODE, the
-Genome in a Bottle project and other sources. Chromatin accessibility and CpG methylation are read from the same
+individuals sequenced by the Human Pangenome Reference Consortium and five from rare disease
+cases. Chromatin accessibility and CpG methylation are read from the same
 molecules in the same experiment, so both are kept in one table here and can be compared without
 worrying about differences in cell preparation or sequencing depth. Six kinds of data are
 available for each sample:
 </p>
 
 <ul>
   <li>Percent accessible: the fraction of Fiber-seq molecules on which a position was called
       accessible, combining both chromosomes.</li>
   <li>FIRE peaks: the accessible regulatory elements called from that signal, with a score and
       a false discovery rate.</li>
   <li>Haplotype accessibility: the percent-accessible signal computed separately for the two
       parental chromosomes and drawn as an overlay, which makes elements that are open on one
       chromosome but not the other visible directly.</li>
   <li>CpG methylation: percent of reads methylated at each CpG, over both chromosomes.</li>
   <li>Haplotype CpG: the same measure computed separately for the two parental chromosomes.</li>
   <li>CpG haplotype difference: the difference in percent methylation between the two
       chromosomes, at four nested significance thresholds.</li>
 </ul>
 
 <p>
 Because 41 samples times six kinds of data is far too many tracks for a checkbox list, samples
 are chosen from a searchable table on this page. Pick the kinds of data you want along the top,
 then select samples in the table; the browser turns on that combination for every sample you
 picked, and keeps each sample's tracks together in the display. Sample class has filter
 checkboxes in the panel to the left, and every column can be searched with the box under its
 heading and clicked to sort, so a sample can be found by name, cell type, sample class or
 accession.
 </p>
 
 <h2>Display Conventions</h2>
 
 <p>
 Accessibility, methylation and the haplotype overlays are all drawn 0 to 100 percent on a fixed
 scale, so heights are comparable between samples and between the two assays. The accessibility
 tracks use maximum as the windowing function, so a narrow element survives zooming out, while
 the methylation tracks use mean, since an average is the meaningful summary for a methylation
 level. FIRE peaks are shown in dense mode by default, one row per sample.
 </p>
 
 <p>
 In both haplotype overlays:
 </p>
 
 <table class="stdTbl">
   <tr><th style="background-color:#0072B2;width:2em">&nbsp;</th><td>Haplotype 1</td></tr>
   <tr><th style="background-color:#D55E00;width:2em">&nbsp;</th><td>Haplotype 2</td></tr>
 </table>
 
 <p>
 The two are overlaid transparently, so a position with equal signal on both chromosomes appears
 as the two colors superimposed, and a haplotype-selective element appears as one color standing
 alone. Which parental chromosome is haplotype 1 is arbitrary and is not consistent between
 samples.
 </p>
 
 <p>
 The CpG haplotype difference track runs from -100 to +100 percent, so a bar above the midline
 means haplotype 1 is more methylated and a bar below it means haplotype 2 is. It is a stack of
 four overlaid signals, one per significance threshold, drawn least significant first so that the
 more significant levels are painted on top:
 </p>
 
 <table class="stdTbl">
   <tr><th style="background-color:#898F8F;width:2em">&nbsp;</th><td>All measured differences, regardless of significance</td></tr>
   <tr><th style="background-color:#EBE534;width:2em">&nbsp;</th><td>p &lt; 0.01</td></tr>
   <tr><th style="background-color:#F59416;width:2em">&nbsp;</th><td>p &lt; 0.001</td></tr>
   <tr><th style="background-color:#FF0000;width:2em">&nbsp;</th><td>p &lt; 0.0001</td></tr>
 </table>
 
 <p>
 The thresholds are nested, so a position drawn red also belongs to all three looser sets.
 Reading the track amounts to reading the color: grey is noise, red is a strong difference
 between the two chromosomes at that CpG.
 </p>
 
 <p>
 The color swatches next to the Sample class filters are:
 </p>
 
 <table class="stdTbl">
   <tr><th style="background-color:#0072B2;width:2em">&nbsp;</th>
       <td>HPRC, a lymphoblastoid (B-lymphocyte, EBV) line from the Human Pangenome Reference
           Consortium</td></tr>
   <tr><th style="background-color:#D55E00;width:2em">&nbsp;</th>
-      <td>Common cell line. Seven of these are lymphoblastoid as well but come from elsewhere:
-          GM12878 from ENCODE, HG002 from the Genome in a Bottle project, and GM25455, GM25456,
-          GM27730, GM28570 and GM28572. The classification is the one supplied by the
-          laboratory, not one inferred from the cell type</td></tr>
+      <td>Common cell line. Two of these are lymphoblastoid as well but come from elsewhere:
+          GM12878 from ENCODE and HG002 from the Genome in a Bottle project</td></tr>
+  <tr><th style="background-color:#009E73;width:2em">&nbsp;</th>
+      <td>Rare disease sample, from a case consented to broad genomic data sharing. Five so far:
+          GM25455, GM25456, GM27730, GM28570 and GM28572</td></tr>
 </table>
 
+<p>
+The classification is the one supplied by the laboratory, not one inferred from the cell type.
+</p>
+
 <p>
 Peaks carry two filterable values, the FIRE score in the signalValue field and the false
 discovery rate as a -log10 value in the qValue field, and both can be filtered from a peak
 track's own configuration page, along with the score. No filter is applied by default. The FDR
 value stops at 100, which is the highest the pipeline reports rather than a real ceiling on
 significance, and 9 percent of the peaks in this track sit at it; filtering at the top of that
 range therefore selects a large group rather than a handful of outstanding peaks. A short
 tick inside each peak marks the point source, the single base the pipeline picked as the summit.
 Switching a peak track to pack or full also gives each peak a mouseover with its FIRE score and
 FDR; dense mode has no per-peak hover, which is a property of dense display rather than of this
 track. The pValue field of the source files is set to -1 throughout and carries no information.
 </p>
 
 <h2>Methods</h2>
 
 <p>
 Permeabilized cells were treated with the Hia5 N6-adenine methyltransferase, which methylates
 adenines in DNA not protected by a bound protein, and high molecular weight DNA was prepared
 into PacBio SMRTbell libraries and sequenced. The adenine methylation added by the enzyme is
 chemically distinct from native CpG methylation, so both are read from the same molecule.
 Adenine methylation was called with fibertools-rs v0.4, and reads were aligned and
 haplotype-phased; for GM12878 an average 20 kb read spans at least one heterozygous variant and
 87.9 percent of reads could be phased against GRCh38.
 </p>
 
 <p>
 The FIRE pipeline v0.0.4, a Snakemake workflow, applied a semi-supervised XGBoost classifier to
 label methyltransferase-sensitive patches on each read as FIRE elements. The classifier was
 trained with Mokapot over 15 iterations on 21 GM12878 experiments spanning 5.8 to 13.3 percent
 adenine methylation, with DNase I and CTCF ChIP-seq peaks as mixed-positive labels. The
 aggregate FIRE score at a position is -50/R times the sum over covering elements of
 log10(1 - min(EP, 0.99)), where R is the read depth and EP the estimated precision of each
 element, which puts the score between 0 and 100; positions covered by fewer than four FIRE
 elements are not scored. The false discovery rate was estimated by shuffling whole reads within
 a chromosome and comparing the resulting score distribution to the observed one. Peaks are FIRE
 score local maxima below a 5 percent FDR with at least 10 percent of covering reads actuated;
 adjacent maxima sharing half their elements or overlapping reciprocally by 90 percent were
 merged, and peak boundaries were set to the median start and end of the underlying elements.
 </p>
 
 <p>
 Base-level CpG methylation was called with jasmine, and the percent methylation at each genomic
 position was computed from a pileup of reads using pb-CpG-tools. Reads were haplotype-phased
 before the pileup, which gives the per-haplotype values, and the difference track is the
 subtraction of one haplotype from the other with the per-position significance thresholds
 applied. See Vollger et al. for the full description of all of the above.
 </p>
 
 <p>
 The bigWig and bigBed files were downloaded from
 <a href="https://s3.kopah.uw.edu/userprod/web/public/hashed.PacBio-Fiber-seq/" target="_blank">the
 Stergachis lab data server</a>, eleven files per sample. The signal files were copied without
 modification. The peak files were rebuilt, because their bigBed header recorded three data
 columns while the data has the full ten of a narrowPeak file, which left the browser unable to
 see the signalValue and qValue columns for filtering or display. The rebuild corrects the header
 and rounds the FIRE score and the FDR to three decimals, which is well beyond the precision
 either measure carries. It does drop peaks called on chrEBV, the Epstein-Barr virus decoy of the
 GRCh38 analysis set, since that sequence is not part of hg38: 309 of 9,253,902 peaks, in 20 of
 the 41 samples, between 2 and 54 peaks each, leaving 9,253,593 in the track. The download,
 rebuild and integrity steps are documented in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/fiberSeq.txt"
 target="_blank">makeDoc</a>, and the scripts that fetch the data and generate the track
 configuration are in the
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/fiberSeq"
 target="_blank">kent source tree</a>.
 </p>
 
 <p>
 GM12878 (accession PM00001) was reprocessed by the lab in September 2026 and every file for that
 sample was replaced here. The earlier release had two placeholder haplotype accessibility bigWigs
 covering a single base, so its haplotype overlay drew nothing. The lab re-ran the sample with
 parental sequence added for phasing, and those tracks now carry full data. Its peak calls and
 CpG haplotype values changed with the reprocessing as well, so figures made from the first
 version of this track will not reproduce exactly for that sample.
 </p>
 
 <h2>Data Access</h2>
 
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to spreadsheet or
 tab-sep tables. From scripts, the data can be accessed through our
 <a href="https://api.genome.ucsc.edu">API</a>, one sample and data type at a time, for example
 track=<i>fiberSeqCompendium_PM00001_acc</i> for GM12878 accessibility or
 <i>fiberSeqCompendium_PM00001_cpg</i> for its CpG methylation. The container name
 <i>fiberSeqCompendium</i> holds no data of its own and cannot be queried, and the peak
 subtracks are not yet available through the API.
 </p>
 
 <p>
 For automated download and analysis, the data are stored in bigWig and bigBed files that can be
 downloaded from <a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/fiberSeq/" target="_blank">our
 download server</a>, one directory per sample accession. Each directory holds the peaks as
 <tt>fire-peaks.ucsc.bb</tt>, the accessibility signal as <tt>all.percent.accessible.bw</tt> with
 <tt>hap1.percent.accessible.bw</tt> and <tt>hap2.percent.accessible.bw</tt>, and the methylation
 as <tt>cpg.combined.bw</tt>, <tt>cpg.hap1.bw</tt>, <tt>cpg.hap2.bw</tt> and four
 <tt>cpg.diffs_*.bw</tt> files. Individual regions or the whole genome annotation can be obtained
 using our tools <tt>bigBedToBed</tt> and <tt>bigWigToBedGraph</tt>, which can be compiled from
 the source code or downloaded as precompiled binaries for your system. Instructions for
 downloading source code and binaries can be found
 <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">here</a>. The tools
 can also be used to obtain features within a given range, e.g.
 <tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/fiberSeq/PM00001/fire-peaks.ucsc.bb
 -chrom=chr21 -start=0 -end=100000000 stdout</tt>
 </p>
 
 <p>
 The mapping from sample accession to sample name and cell type is in
 <a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/fiberSeq/fiberSeqCompendium_metadata.tsv"
 target="_blank">fiberSeqCompendium_metadata.tsv</a>. The original data can be downloaded from the
 <a href="https://s3.kopah.uw.edu/userprod/web/public/hashed.PacBio-Fiber-seq/" target="_blank">Stergachis
 lab data server</a>, and the lab maintains its own track hub and documentation at
 <a href="https://fiberseq.github.io/" target="_blank">fiberseq.github.io</a>. The FIRE pipeline
 is at <a href="https://github.com/fiberseq/FIRE" target="_blank">github.com/fiberseq/FIRE</a>, the
 adenine methylation caller at
 <a href="https://github.com/fiberseq/fibertools-rs" target="_blank">github.com/fiberseq/fibertools-rs</a>,
 and the CpG pileup tool at
 <a href="https://github.com/PacificBiosciences/pb-CpG-tools" target="_blank">github.com/PacificBiosciences/pb-CpG-tools</a>.
 </p>
 
 <h2>Credits</h2>
 
 <p>
 Thanks to Mitchell Vollger, Andrew Stergachis and Shane Neph for generating this data, for
 assembling it into track hubs and for their help in arranging these tracks for the browser.
 </p>
 
 <h2>References</h2>
 
 <p>
 Vollger MR, Swanson EG, Neph SJ, Ranchalis J, Munson KM, Ho CH, Cheng YHH, Sede&#241;o-Cort&#233;s AE, Fondrie
 WE, Bohaczuk SC <em>et al</em>.
 <a href="https://doi.org/10.1101/2024.06.14.599122" target="_blank">
 A haplotype-resolved view of human gene regulation</a>.
 <em>bioRxiv</em>. 2025 Jun 2;.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/40501892" target="_blank">40501892</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12157683/" target="_blank">PMC12157683</a>
 </p>
 
 <p>
 Stergachis AB, Debo BM, Haugen E, Churchman LS, Stamatoyannopoulos JA.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/32587015" target="_blank">
 Single-molecule regulatory architectures captured by chromatin fiber sequencing</a>.
 <em>Science</em>. 2020 Jun 26;368(6498):1449-1454.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/32587015" target="_blank">32587015</a>
 </p>