d0ddd717713f275dc2f567c9b6d733fd3f55bd90 markd Fri Oct 2 21:27:25 2026 -0700 Put PRO-cap and ProCapNet inside a Transcription Initiation folder. refs #35528 The two data sources were top-level superTracks because a superTrack could not be a member of another one. #38460 fixes that, so they now sit inside a transcriptionStart superTrack, which is what the collection was meant to be. The container is written before either of its children, and both children before any of their members: tdbQuery -strict rejects a file in which another track comes between a superTrack and one of its children. This depends on the #38460 build. A plain make here with the installed binaries produces a half-built state, since hgTrackDb drops the container from the table and trackDbToTxt then writes an hs1 curated hub whose member names a parent stanza that is not there. Both makedocs say so. diff --git src/hg/makeDb/doc/hg38/transcriptionStart.txt src/hg/makeDb/doc/hg38/transcriptionStart.txt index f731063c44b..8e8bfc5f548 100644 --- src/hg/makeDb/doc/hg38/transcriptionStart.txt +++ src/hg/makeDb/doc/hg38/transcriptionStart.txt @@ -1,182 +1,185 @@ # Transcription start sites: ENCODE 4 PRO-cap and ProCapNet, #35528, Claude Thu Sep 17 2026 (Claude/markd) # Layout: one superTrack per data source, each holding its multiWig overlays # directly, the way fantom5 does. Not a composite: under a composite hgTrackUi # lists descendant leaves rather than containers, so an overlay gets no # configuration block of its own and hiding both strands of a cell line leaves an # empty row where the overlay was (#38441). As a superTrack member each overlay # is a track in its own right, with a full configuration page, and hiding it # hides the whole overlay. The cost is the subtrack matrix and the sample class # filter, neither of which a superTrack offers. # -# These two superTracks are top level rather than sitting inside a single -# transcriptionStart folder, because superTracks do not nest: a superTrack given -# a parent passes tdbQuery -check -strict and is then silently dropped at load -# (#38460). transcriptionStart.html is left in the tree unused against that -# being fixed. +# The two data-source superTracks sit inside a transcriptionStart superTrack. +# That needs the nesting fix in #38460: before it, a superTrack given a parent +# passed tdbQuery -check -strict and was then silently dropped at load. Until +# that fix is released, a plain make here produces a half-built state, because +# the trackDb make calls two installed binaries that predate it: hgTrackDb drops +# the container from the table, and trackDbToTxt then writes an hs1 curated hub +# whose member names a parent stanza that is not there. Load with a build that +# has the fix, and rebuild the hub, until it ships. # Scripts are in ~/kent/src/hg/makeDb/outside/proCapNet. The cell lines, their # ENCODE accessions and their colors are in proCapNetExperiments.tsv there. ############################################################################## # ProCapNet predictions and sequence-contribution scores ############################################################################## # The predictions are 24 GB per file and are re-encoded, so they are fetched to # a working area outside the track directory. The contribution scores are used # unaltered and go straight to their final home. mkdir -p /hive/data/outside/proCapNet/hg38 /hive/data/genomes/hg38/bed/proCapNet cd /hive/data/genomes/hg38/bed/proCapNet ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetDownload hg38 \ ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv \ /hive/data/outside/proCapNet/hg38/pred contrib # 290 GB, about 25 minutes at six parallel streams # Re-encode the predictions from one bedGraph interval per base into fixedStep # sections, which cuts the size by about a third without changing a value, and # drop the bases holding a literal NaN. Those NaNs are all in unresolved (N) # reference sequence and make the track's autoScale and summary statistics NaN. ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredBuild \ /hive/data/genomes/hg38/chrom.sizes /hive/data/outside/proCapNet/hg38/pred pred 6 # all 12 files report the same counts: # 24 chroms, 3088269832 bases, 2924001250 with data, # 164268582 dropped as no-data or NaN (5.3%) # That is more than the 150610728 gap bases on the primary chromosomes, because # no prediction was made where most of a 2114 bp window was unresolved, which # also takes out sequence flanking each gap. # about 24 GB in, 15.7 to 15.9 GB out per file # proCapNetPredBuild writes the minus strand negated, so it draws below the # baseline from its own values. Doing it in the data rather than with the # trackDb negateValues setting is what makes the composite's own negate control # work: that control sets one shared value for the composite, which replaced the # per-track negateValues and sent both strands the same way, with no way back to # the default short of a cart reset. encode4ProCap never had the problem, # because ENCODE already publishes its minus strand negative. # confirm the summary statistics are numbers again, they were nan before bigWigInfo pred/K562.proCapNet.pos.bw | egrep 'basesCovered|mean|std' # basesCovered: 2,924,001,250 # mean: 0.019478 # std: 0.300502 # spot check that the re-encoding is lossless, comparing random windows of the # converted file against the download python3 - <<'PYEOF' import numpy as np, pyBigWig, random src = pyBigWig.open("/hive/data/outside/proCapNet/hg38/pred/K562.pos.bigWig") out = pyBigWig.open("pred/K562.proCapNet.pos.bw") random.seed(7) bad = 0 for chrom in ("chr1", "chr7", "chr14", "chr21", "chrX"): for _ in range(3): s = random.randrange(0, src.chroms()[chrom] - 200000) a = src.values(chrom, s, s + 200000, numpy=True) b = out.values(chrom, s, s + 200000, numpy=True) na, nb = np.isnan(a), np.isnan(b) if not (na == nb).all() or not (a[~na] == b[~nb]).all(): bad += 1 print("mismatched windows:", bad) PYEOF # mismatched windows: 0 ############################################################################## # negating the minus strand, after the fact ############################################################################## # The prediction files were first built with the minus strand positive and # flipped at display time with the trackDb negateValues setting. That broke the # composite's negate control, as described above, so the twelve minus-strand # files were rewritten with negated values rather than rebuilt from the # downloads, which had already been deleted: # # for db in hg38 hs1 ; do # base=/hive/data/genomes/$db/bed/proCapNet # mkdir -p $base/previousPred # ls $base/pred/*.neg.bw | while read f ; do # ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredToFixedStep \ # --negate /hive/data/genomes/$db/chrom.sizes $f $base/negated/$(basename $f) # done # done # # Each output was checked against its input: same nBasesCovered, value range # mirrored, sampled values exactly negated. The originals were then moved to # previousPred/ and the negated files put in their place, so /gbdb needs no new # symlinks. previousPred/ is about 197 GB and can be removed once the track has # been through QA. A rebuild from the downloads does not need any of this: # proCapNetPredBuild negates the minus strand itself. ############################################################################## # ENCODE 4 PRO-cap ############################################################################## mkdir -p /hive/data/genomes/hg38/bed/encode4ProCap cd /hive/data/genomes/hg38/bed/encode4ProCap # Resolve the portal metadata: for each of the six experiments take only the # plus and minus strand signal of unique reads files belonging to that # experiment's default analysis, which drops files superseded by a later # reprocessing. ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetEncodeMeta \ ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv encodeFiles.tsv # A673 ENCSR046BCI 4 files # Caco-2 ENCSR100LIJ 4 files # Calu3 ENCSR935RNW 4 files # HUVEC ENCSR098LLB 2 files # K562 ENCSR261KBX 4 files # MCF10A ENCSR799DGV 8 files # A copy of the resulting file is committed as # outside/proCapNet/proCapNetEncodeFiles.tsv. It is the record of which ENCODE # file accessions went into each track: re-running the step later can resolve to # different files, because the default analysis of an experiment changes when # ENCODE reprocesses it. # ENCODE publishes no pooled PRO-cap file, so download the per-replicate files # and sum them per cell line and strand. MCF10A has two sequencing runs per # biological replicate, hence eight files. Minus strand signal is negative as # released and is kept that way. ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetEncodeBuild \ encodeFiles.tsv /hive/data/genomes/hg38/chrom.sizes \ /hive/data/outside/proCapNet/encode/raw . # K562.ENCSR261KBX.pos.bw: 2 inputs, 3227685 input intervals, 3051479 merged intervals, # 3194212 bases covered # every input interval is accounted for; total signal is conserved exactly # sanity check: the summed signal equals the sum of the inputs python3 - <<'PYEOF' import pyBigWig def total(f): bw = pyBigWig.open(f); s = bw.header()["sumData"]; bw.close(); return s raw = "/hive/data/outside/proCapNet/encode/raw" print(total(raw + "/ENCFF994CSC.pos.bigWig") + total(raw + "/ENCFF580QEK.pos.bigWig"), total("K562.ENCSR261KBX.pos.bw")) PYEOF # 9605200.0 9605200.0 ############################################################################## # trackDb ############################################################################## # The stanzas are generated. These are traditional composites: a faceted one is # routed to facetedCompositeUi(), which returns before the wiggle controls are # drawn, and signal tracks need those. The matrix is declared over the leaf # bigWigs, so each strand is a cell; the multiWig containers carry no subGroups # because a container spans both strands. cd ~/kent/src/hg/makeDb/outside/proCapNet ./proCapNetTrackDb hg38 proCapNetExperiments.tsv \ ~/kent/src/hg/makeDb/trackDb/human/hg38/transcriptionStart.ra # transcriptionStart.ra is generated; the include line in human/hg38/trackDb.ra # is added by hand: # include transcriptionStart.ra alpha ############################################################################## # cleanup ############################################################################## # The downloads are only needed to build the tracks: the predictions for the # re-encoding, the ENCODE per-replicate files for the summing. /hive charges # twice the apparent size, so this returns about 580 GB of disk. rm -rf /hive/data/outside/proCapNet/hg38/pred /hive/data/outside/proCapNet/encode