156d289f517d82c4d5a8c59981dfec21a715af8d
lrnassar
  Mon Sep 21 09:22:45 2026 -0700
Fiber-seq QA fixes: take sample class from the lab's own table, pin the overlay
draw order, and correct the description pages. refs #36210

Sample class had been guessed from the free-text cell type with a two-name
exception list, which filed five lymphoblastoid lines as HPRC that are not.
It now comes from a fifth column in fiberSeqSamples.tsv carrying the
classification the lab supplies, and sampleClass()/NOT_HPRC are gone. Facet
counts go from 25/16 to 20/21.

The CpG difference track is a solidOverlay whose four files hold the same value
at a shared base, so the tier drawn last is the colour the reader sees. Its
children inherited a single priority from the parent, trackPriCmp ties on that,
and slSort is not stable, so the paint order was arbitrary and p<0.01 was
covering p<0.0001. The four levels and the two haplotype children now carry
explicit priorities. At chr20:29,300,000-29,305,000 red goes from 2 image
columns to 88.

CpG haplotype children take their shortLabel prefix from their container, so
they read "<sample> CpG Hap1" rather than colliding with the accessibility
children's "<sample> Hap1". 82 labels were duplicated.

Description pages: fiberSeqAcc.html attached the FIRE score's "fewer than four
elements" cutoff to the percent accessible signal, which is a different
quantity; two pages overstated how many of the lymphoblastoid lines come from
HPRC; both pages told readers to query the API with container track names,
which it refuses by design. Also documents GM12878's trio phasing and the FDR
ceiling at 100, corrects the difference track's stated range, replaces a
non-ASCII author name with numeric entities, and adds db= to hgTrackUi links.

makeDoc: corrects a basesCovered figure that was out by a factor of a hundred,
refreshes the peak file sizes after the PM00001 reissue, rewrites the
sampleClass rationale, and records why multiWig children need their own
priority.

diff --git src/hg/makeDb/doc/hg38/fiberSeq.txt src/hg/makeDb/doc/hg38/fiberSeq.txt
index 4a2bc77533c..192c65cc397 100644
--- src/hg/makeDb/doc/hg38/fiberSeq.txt
+++ src/hg/makeDb/doc/hg38/fiberSeq.txt
@@ -69,32 +69,33 @@
 ~/kent/src/hg/makeDb/scripts/fiberSeq/fiberSeqDownload.sh \
     /hive/data/genomes/hg38/bed/fiberSeq 12 > download.log 2>&1
 
 grep -c '^got' download.log
 grep -c FAIL download.log
 
 # Check every file parses and holds data:
 ~/kent/src/hg/makeDb/scripts/fiberSeq/fiberSeqCheck.sh /hive/data/genomes/hg38/bed/fiberSeq
 
 # On the first mirror this reported two placeholder files at the source:
 #   PM00001/hap1.percent.accessible.bw  (512 bytes, basesCovered 1)
 #   PM00001/hap2.percent.accessible.bw  (512 bytes, basesCovered 1)
 # They were valid bigWigs covering a single base with a value of zero, so
 # GM12878's haplotype accessibility overlay drew nothing.  Note that they were
 # NOT basesCovered 0, which is why the check tests for less than 1000; the
-# smallest legitimate haplotype file here covers 4.1 million bases (PM00008,
-# the near-haploid Hap1 line, which has little phaseable heterozygosity).
+# smallest legitimate haplotype file here covers 399,383,234 bases (PM00008
+# hap1, the near-haploid Hap1 line, which has little phaseable heterozygosity;
+# every other sample's haplotype files cover 2.1 billion bases or more).
 # Reported to the lab, and fixed by them - see the PM00001 section below.
 
 # ---------------------------------------------------------------------------
 # 2026-09-14  PM00001 (GM12878) reissued
 # ---------------------------------------------------------------------------
 
 # The lab reprocessed GM12878 and replaced the files IN PLACE, under the same
 # per-sample hash directory, so fiberSeqSamples.tsv did not change.  Ten of its
 # twelve files differ from the first mirror; only cpg.combined.bw is unchanged.
 # A size sweep over all 41 samples confirmed PM00001 is the only one affected.
 
 ~/kent/src/hg/makeDb/scripts/fiberSeq/fiberSeqDownload.sh \
     /hive/data/genomes/hg38/bed/fiberSeq 10 PM00001 > download.PM00001.log 2>&1
 
 # The two haplotype files are now real: basesCovered 2,894,029,052 (hap1) and
@@ -185,32 +186,32 @@
 #
 # Reconciliation, from peakCounts.tsv (accession, source peaks, peaks in the
 # track, peaks dropped on chrEBV):
 #   source peaks     9,253,902
 #   peaks in track   9,253,593
 #   dropped chrEBV         309   in 20 of the 41 samples, 2 to 54 each
 # Every sample balances exactly: source - chrEBV == in track.  The largest
 # single loss is GM12878 with 54 peaks, which is 0.03 percent of its 196,742.
 # This is mentioned in the methods section of fiberSeqCompendium.html.
 # (Totals as of the PM00001 reissue; before it they were 9,487,043 source,
 # 9,486,622 in track and 421 on chrEBV, all of the difference in that sample.)
 
 # The rebuild also rounds signalValue and qValue to 3 decimals.  They arrive
 # with full double precision (qValues like 22.807437495448326), which reads
 # badly in a mouseover and on the details page and costs a third of the file
-# size: 467 MB over the 41 files before rounding, 313 MB after.  pValue is the
-# sentinel -1 throughout and is untouched.
+# size: 500 MB over the 41 files before rounding, 320 MB after (recounted after
+# the PM00001 reissue).  pValue is the sentinel -1 throughout and is untouched.
 #
 # The rebuilt files are named fire-peaks.ucsc.bb and are what trackDb points
 # at; the originals are kept next to them for comparison.
 
 # Reconciliation is re-derived by the script on every run, into peakFixReport.txt:
 sed -E 's/^([A-Z0-9]+) ([0-9]+) peaks \(source ([0-9]+), dropped ([0-9]+) on chrEBV\)$/\1\t\2\t\3\t\4/' \
     peakFixReport.txt | awk -F'\t' '{kept+=$2; src+=$3; ebv+=$4; if($3-$4!=$2) mism++}
     END{printf "kept %d source %d chrEBV %d mismatched %d\n", kept, src, ebv, mism+0}'
 # kept 9253593 source 9253902 chrEBV 309 mismatched 0
 #
 # fixOne() skips a sample whose fire-peaks.ucsc.bb is newer than its source, so
 # after the PM00001 reissue only that one rebuilt and the other 40 said "have".
 # That also means the run only prints one "got" line, and peakFixReport.txt has
 # to have PM00001's row replaced rather than being rewritten wholesale.
 
@@ -296,41 +297,65 @@
 #    bigBed-generic filter.<field> settings are NOT read by this type.  All
 #    three defaults are the full range, so nothing is hidden until the user
 #    narrows one.  Verified in the rendered UI:
 #      hgTrackUi?db=hg38&g=fiberSeqCompendium_PM00004_peaks
 #    draws min/max boxes for all three with the right limit hints.
 #
 # 5. Facet columns: only "sampleClass" is faceted.  facetedComposite.js offers
 #    a facet value only when it occurs more than once (a checkbox matching a
 #    single row is just a slow search box), so:
 #      accession  41 distinct, all count 1, and excluded anyway as primaryKey
 #      _sample    41 distinct, all count 1 - can never be a facet
 #      _cellType  14 distinct but only 2 with count > 1 (Lymphoblastoid 27,
 #                 Embryonic stem cell 2), so as a facet it drew two checkboxes
 #                 and left 12 samples unreachable.  Underscored, so it is a
 #                 searchable and sortable column instead.
-#      sampleClass 3 values, all count > 1.  Derived from the lab's free-text
-#                 cell type, and worth replacing when they give us real HPRC
-#                 metadata (donor sex, population) that would facet properly.
+#      sampleClass 2 values, HPRC (20) and Common Cell Line (21), both count > 1.
+#                 Read from the fifth column of fiberSeqSamples.tsv, which
+#                 carries the lab's own classification (the Note column of
+#                 Mitchell's sample sheet).  It used to be derived from the
+#                 free-text cell type, which misfiled five lymphoblastoid lines
+#                 that are not HPRC - see the note below.  Still worth extending
+#                 when the lab sends real donor metadata (sex, population), which
+#                 would give more than one useful facet.
 #
 # 6. Subtracks need an explicit priority.  Without one they fall back to a label
 #    sort, which showed a sample's data types as Peaks, CpG, Acc on a first
 #    visit.  The script now numbers them sample-outer, declared-data-type-inner
 #    (i*10 + j + 1), which matches the row of data type checkboxes across the
 #    top of the table.  cartDump.c clears "<mdid>_*.priority" and writes its own
 #    on every submit, so this only sets the starting order.
 #
+#    The children INSIDE each multiWig need their own priority too, and there it
+#    is not cosmetic.  makeContainerTrack() in hg/hgTracks/container.c sorts a
+#    container's children with trackPriCmp, which compares priority alone and
+#    returns 0 on a tie, and slSort is not stable - so children left on the
+#    parent's inherited priority draw in an arbitrary order.  For cpgDiff that
+#    decides the picture: it is a solidOverlay and all four files hold the same
+#    value at a shared base, so the level drawn last is the colour the user sees.
+#    Left unset, p<0.01 painted over p<0.001 and p<0.0001, i.e. the least
+#    stringent tier won and the track said the opposite of what it meant.
+#    Measured at chr20:29,300,000-29,305,000, where 107 of the 133 p<0.01 bases
+#    are also p<0.0001: before, 98 image columns showed only yellow and 2 showed
+#    red; after giving the four levels priority 1-4, 88 columns show red and 51
+#    show grey.  The two haplotype children carry priority 1 and 2 for the same
+#    reason, though transparentOverlay makes their order much less visible.
+#
+#    The haplotype children also take their shortLabel prefix from the
+#    container's own shortLabel, so the CpG pair reads "<sample> CpG Hap1/2"
+#    rather than colliding with the accessibility pair's "<sample> Hap1/2".
+#
 # 7. mouseOver text only appears in pack or full.  Dense makes no per-item map
 #    boxes at all, so no bigBed-like track has a per-item hover there.  Since
 #    Andrew asked for peaks in dense by default, the mouseover is there for
 #    whoever switches a peak track to pack.  Checked by temporarily dropping
 #    onlyVisibility and reading the image map: in pack the AREA tags carry
 #      data-tooltip='<b>K562 FIRE peak</b><br>FIRE score: 17.76015
 #                    <br>-log10 FDR: 22.807437495448326<br>Score: 177'
 #    and in dense there are no per-peak AREA tags.  Note the tooltip is HTML
 #    entity encoded in the page, so grep for it with the entities in mind.
 
 # ---------------------------------------------------------------------------
 # Rendering checks
 # ---------------------------------------------------------------------------
 
 # ACTB promoter, the accessibility overlay, all seven cell lines in color: