54f465e316441236c128491e53c0e5464c0e672c lrnassar Tue Sep 29 11:40:49 2026 -0700 Fiber-seq: bring the fiberSeqTrackDb.py docstrings back in line with the three-value sampleClass, and stop the Compendium intro from keeping a running tally of where the lymphoblastoid lines came from. readSamples still said the five GM lines were common cell lines and writeMetadata still said there were two classes, both a few lines from the SAMPLE_CLASS_COLORS entry that added the third. writeMetadata also still had "Common Cell Line" in the old Title Case, which yesterday's rename missed because the string is wrapped across two source lines and a line oriented sed cannot see it; worth remembering for the next rename. The intro sentence had grown a breakdown that did not add up, 20 HPRC plus 5 rare disease against 27 lymphoblastoid lines, leaving GM12878 and HG002 unaccounted; it now points at the Sample class filter instead of counting, since the lab has said more rare disease samples are coming and the tally would go stale again. While in there, the claim that accession order keeps each class together is softened to what the data actually does: the common cell lines fall in two runs either side of the HPRC block. Generated output is unchanged; the .ra, the metadata TSV and the colors JSON all regenerate byte identical. Caught by Claude review of 10f0f6d516. refs #36210 diff --git src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html index 3a0b84c7f99..56947a9046b 100644 --- src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html +++ src/hg/makeDb/trackDb/human/hg38/fiberSeqCompendium.html @@ -1,249 +1,249 @@

Description

This track is part of the Fiber-seq collection. It holds the -full Fiber-seq data for 41 samples: 14 cell lines and 27 lymphoblastoid lines, 20 of them from -individuals sequenced by the Human Pangenome Reference Consortium and five from rare disease -cases. Chromatin accessibility and CpG methylation are read from the same +full Fiber-seq data for 41 samples: 14 cell lines and 27 lymphoblastoid lines. The Sample class +filter below splits them by where they came from. Chromatin accessibility and CpG methylation +are read from the same molecules in the same experiment, so both are kept in one table here and can be compared without worrying about differences in cell preparation or sequencing depth. Six kinds of data are available for each sample:

Samples are chosen from the searchable table on this page. Pick the kinds of data you want along the top, then select samples in the table; the browser turns on that combination for every sample you picked, and keeps each sample's tracks together in the display. Sample class has filter checkboxes in the panel to the left, and every column can be searched with the box under its heading and clicked to sort, so a sample can be found by name, cell type, sample class or accession.

Display Conventions

Accessibility, methylation and the haplotype overlays are all drawn 0 to 100 percent on a fixed scale, so heights are comparable between samples and between the two assays. The accessibility tracks use maximum as the windowing function, so a narrow element survives zooming out, while the methylation tracks use mean, since an average is the meaningful summary for a methylation level. FIRE peaks are shown in dense mode by default, one row per sample.

In both haplotype overlays:

 Haplotype 1
 Haplotype 2

The two are overlaid transparently, so a position with equal signal on both chromosomes appears as the two colors superimposed, and a haplotype-selective element appears as one color standing alone. Which parental chromosome is haplotype 1 is arbitrary and is not consistent between samples.

The CpG haplotype difference track runs from -100 to +100 percent, so a bar above the midline means haplotype 1 is more methylated and a bar below it means haplotype 2 is. Each position is colored by the most stringent significance threshold it meets:

 All measured differences, regardless of significance
 p < 0.01
 p < 0.001
 p < 0.0001

The thresholds are nested, so a position drawn red also belongs to all three looser sets. Reading the track amounts to reading the color: grey is noise, red is a strong difference between the two chromosomes at that CpG.

The color swatches next to the Sample class filters are:

  HPRC, a lymphoblastoid (B-lymphocyte, EBV) line from the Human Pangenome Reference Consortium
  Common cell line. Two of these are lymphoblastoid as well but come from elsewhere: GM12878 from ENCODE and HG002 from the Genome in a Bottle project
  Rare disease sample, from a case consented to broad genomic data sharing. Five so far: GM25455, GM25456, GM27730, GM28570 and GM28572

The classification is the one supplied by the laboratory, not one inferred from the cell type.

Peaks carry two filterable values, the FIRE score in the signalValue field and the false discovery rate as a -log10 value in the qValue field, and both can be filtered from a peak track's own configuration page, along with the score. No filter is applied by default. The FDR value stops at 100, which is the highest the pipeline reports rather than a real ceiling on significance, and 9 percent of the peaks in this track sit at it; filtering at the top of that range therefore selects a large group rather than a handful of outstanding peaks. A short tick inside each peak marks the point source, the single base the pipeline picked as the summit. Switching a peak track to pack or full gives each peak a mouseover with its FIRE score and FDR. The pValue field of the source files is set to -1 throughout and carries no information.

Methods

Permeabilized cells were treated with the Hia5 N6-adenine methyltransferase, which methylates adenines in DNA not protected by a bound protein, and high molecular weight DNA was prepared into PacBio SMRTbell libraries and sequenced. The adenine methylation added by the enzyme is chemically distinct from native CpG methylation, so both are read from the same molecule. Adenine methylation was called with fibertools-rs v0.4, and reads were aligned and haplotype-phased; for GM12878 an average 20 kb read spans at least one heterozygous variant and 87.9 percent of reads could be phased against GRCh38.

The FIRE pipeline v0.0.4, a Snakemake workflow, applied a semi-supervised XGBoost classifier to label methyltransferase-sensitive patches on each read as FIRE elements. The classifier was trained with Mokapot over 15 iterations on 21 GM12878 experiments spanning 5.8 to 13.3 percent adenine methylation, with DNase I and CTCF ChIP-seq peaks as mixed-positive labels. The aggregate FIRE score at a position is -50/R times the sum over covering elements of log10(1 - min(EP, 0.99)), where R is the read depth and EP the estimated precision of each element, which puts the score between 0 and 100; positions covered by fewer than four FIRE elements are not scored. The false discovery rate was estimated by shuffling whole reads within a chromosome and comparing the resulting score distribution to the observed one. Peaks are FIRE score local maxima below a 5 percent FDR with at least 10 percent of covering reads actuated; adjacent maxima sharing half their elements or overlapping reciprocally by 90 percent were merged, and peak boundaries were set to the median start and end of the underlying elements.

Base-level CpG methylation was called with jasmine, and the percent methylation at each genomic position was computed from a pileup of reads using pb-CpG-tools. Reads were haplotype-phased before the pileup, which gives the per-haplotype values, and the difference track is the subtraction of one haplotype from the other with the per-position significance thresholds applied. See Vollger et al. for the full description of all of the above.

The bigWig and bigBed files were downloaded from the Stergachis lab data server, eleven files per sample. The signal files were copied without modification. The peak files were rebuilt, because their bigBed header recorded three data columns while the data has the full ten of a narrowPeak file, which left the browser unable to see the signalValue and qValue columns for filtering or display. The rebuild corrects the header and rounds the FIRE score and the FDR to three decimals, which is well beyond the precision either measure carries. It does drop peaks called on chrEBV, the Epstein-Barr virus decoy of the GRCh38 analysis set, since that sequence is not part of hg38: 309 of 9,253,902 peaks, in 20 of the 41 samples, between 2 and 54 peaks each, leaving 9,253,593 in the track. The download, rebuild and integrity steps are documented in the makeDoc, and the scripts that fetch the data and generate the track configuration are in the kent source tree.

GM12878 (accession PM00001) was reprocessed by the lab in September 2026 and every file for that sample was replaced here. The earlier release had two placeholder haplotype accessibility bigWigs covering a single base, so its haplotype overlay drew nothing. The lab re-ran the sample with parental sequence added for phasing, and those tracks now carry full data. Its peak calls and CpG haplotype values changed with the reprocessing as well, so figures made from the first version of this track will not reproduce exactly for that sample.

Data Access

The data can be explored interactively in table format with the Table Browser or the Data Integrator and exported from there to spreadsheet or tab-sep tables. From scripts, the data can be accessed through our API, one sample and data type at a time, for example track=fiberSeqCompendium_PM00001_acc for GM12878 accessibility or fiberSeqCompendium_PM00001_cpg for its CpG methylation. The collection name fiberSeqCompendium holds no data of its own and cannot be queried. The FIRE peaks are not yet available through the API; use the Table Browser or the downloadable files for those.

For automated download and analysis, the data are stored in bigWig and bigBed files that can be downloaded from our download server, one directory per sample accession. Each directory holds the peaks as fire-peaks.ucsc.bb, the accessibility signal as all.percent.accessible.bw with hap1.percent.accessible.bw and hap2.percent.accessible.bw, and the methylation as cpg.combined.bw, cpg.hap1.bw, cpg.hap2.bw and four cpg.diffs_*.bw files. Individual regions or the whole genome annotation can be obtained using our tools bigBedToBed and bigWigToBedGraph, which can be compiled from the source code or downloaded as precompiled binaries for your system. Instructions for downloading source code and binaries can be found here. The tools can also be used to obtain features within a given range, e.g. bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/fiberSeq/PM00001/fire-peaks.ucsc.bb -chrom=chr21 -start=0 -end=100000000 stdout

The mapping from sample accession to sample name and cell type is in fiberSeqCompendium_metadata.tsv. The original data can be downloaded from the Stergachis lab data server, and the lab maintains its own track hub and documentation at fiberseq.github.io. The FIRE pipeline is at github.com/fiberseq/FIRE, the adenine methylation caller at github.com/fiberseq/fibertools-rs, and the CpG pileup tool at github.com/PacificBiosciences/pb-CpG-tools.

Credits

Thanks to Mitchell Vollger, Andrew Stergachis and Shane Neph for generating this data, for assembling it into track hubs and for their help in arranging these tracks for the browser.

References

Vollger MR, Swanson EG, Neph SJ, Ranchalis J, Munson KM, Ho CH, Cheng YHH, Sedeño-Cortés AE, Fondrie WE, Bohaczuk SC et al. A haplotype-resolved view of human gene regulation. bioRxiv. 2025 Jun 2;. PMID: 40501892; PMC: PMC12157683

Stergachis AB, Debo BM, Haugen E, Churchman LS, Stamatoyannopoulos JA. Single-molecule regulatory architectures captured by chromatin fiber sequencing. Science. 2020 Jun 26;368(6498):1449-1454. PMID: 32587015