0c42aea56b751a2b4a58feb225a9659819a90994 markd Sat Oct 3 06:22:01 2026 -0700 Nest the TSS tracks on alpha only, leaving the release as it is. refs #35528 PRO-cap and ProCapNet sit inside a transcriptionStart superTrack on alpha. That needs #38460, which is on master but not in a release, and these tracks are already released, so each source is written twice with complementary release tags: the nested copy alpha, the flat copy beta,public. hgTrackDb takes the copy whose release matches and rejects two whose releases overlap. Verified that beta and public are untouched: a public build from this file is byte-identical to a public build from the released flat tree, same md5 over tableName, shortLabel, type, visibility, priority and settings. tdbQuery -check -strict passes on all three releases, and an alpha build has the container with both members parented while beta and public have neither. Drop the release tags and the flat copies once #38460 ships. The makedocs also record that the alpha build needs hgTrackDb and trackDbToTxt in /cluster/bin/x86_64 to carry the #38460 fix. make alpha does not install there, BINDIR defaults to ~/bin/$MACHTYPE, so they were put in by hand and the weekly utils build will overwrite them. diff --git src/hg/makeDb/doc/hg38/transcriptionStart.txt src/hg/makeDb/doc/hg38/transcriptionStart.txt index 5174803edd9..83b1a762bdf 100644 --- src/hg/makeDb/doc/hg38/transcriptionStart.txt +++ src/hg/makeDb/doc/hg38/transcriptionStart.txt @@ -1,184 +1,194 @@ # Transcription start sites: ENCODE 4 PRO-cap and ProCapNet, #35528, Claude Thu Sep 17 2026 (Claude/markd) # Layout: one superTrack per data source, each holding its multiWig overlays # directly, the way fantom5 does. Not a composite: under a composite hgTrackUi # lists descendant leaves rather than containers, so an overlay gets no # configuration block of its own and hiding both strands of a cell line leaves an # empty row where the overlay was (#38441). As a superTrack member each overlay # is a track in its own right, with a full configuration page, and hiding it # hides the whole overlay. The cost is the subtrack matrix and the sample class # filter, neither of which a superTrack offers. # -# These two superTracks are top level rather than sitting inside a single -# transcriptionStart folder, because superTracks do not nest: a superTrack given -# a parent passes tdbQuery -check -strict and is then silently dropped at load -# (#38460). transcriptionStart.html is left in the tree unused against that -# being fixed. +# On alpha the two data-source superTracks sit inside a transcriptionStart +# superTrack; on beta and public they stand on their own. Nesting a superTrack +# inside a superTrack needs the fix in #38460, which is on master but not in a +# release, and these tracks are already released, so each source is written +# twice with complementary release tags: the nested copy alpha, the flat copy +# beta,public. hgTrackDb takes the copy whose release matches and rejects two +# whose releases overlap, so the tags have to stay complementary. The public +# build is byte-identical to what was released. Drop the release tags and the +# flat copies once #38460 ships. +# +# The alpha build here also needs hgTrackDb and trackDbToTxt in +# /cluster/bin/x86_64 to carry the #38460 fix. make alpha does not put anything +# there: BINDIR defaults to ~/bin/$MACHTYPE, and that directory is updated by +# the weekly utils build. They were installed by hand on 2026-10-03 and the +# next weekly build will overwrite them with whatever is in the branch then. # Scripts are in ~/kent/src/hg/makeDb/outside/proCapNet. The cell lines, their # ENCODE accessions and their colors are in proCapNetExperiments.tsv there. ############################################################################## # ProCapNet predictions and sequence-contribution scores ############################################################################## # The predictions are 24 GB per file and are re-encoded, so they are fetched to # a working area outside the track directory. The contribution scores are used # unaltered and go straight to their final home. mkdir -p /hive/data/outside/proCapNet/hg38 /hive/data/genomes/hg38/bed/proCapNet cd /hive/data/genomes/hg38/bed/proCapNet ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetDownload hg38 \ ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv \ /hive/data/outside/proCapNet/hg38/pred contrib # 290 GB, about 25 minutes at six parallel streams # Re-encode the predictions from one bedGraph interval per base into fixedStep # sections, which cuts the size by about a third without changing a value, and # drop the bases holding a literal NaN. Those NaNs are all in unresolved (N) # reference sequence and make the track's autoScale and summary statistics NaN. ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredBuild \ /hive/data/genomes/hg38/chrom.sizes /hive/data/outside/proCapNet/hg38/pred pred 6 # all 12 files report the same counts: # 24 chroms, 3088269832 bases, 2924001250 with data, # 164268582 dropped as no-data or NaN (5.3%) # That is more than the 150610728 gap bases on the primary chromosomes, because # no prediction was made where most of a 2114 bp window was unresolved, which # also takes out sequence flanking each gap. # about 24 GB in, 15.7 to 15.9 GB out per file # proCapNetPredBuild writes the minus strand negated, so it draws below the # baseline from its own values. Doing it in the data rather than with the # trackDb negateValues setting is what makes the composite's own negate control # work: that control sets one shared value for the composite, which replaced the # per-track negateValues and sent both strands the same way, with no way back to # the default short of a cart reset. encode4ProCap never had the problem, # because ENCODE already publishes its minus strand negative. # confirm the summary statistics are numbers again, they were nan before bigWigInfo pred/K562.proCapNet.pos.bw | egrep 'basesCovered|mean|std' # basesCovered: 2,924,001,250 # mean: 0.019478 # std: 0.300502 # spot check that the re-encoding is lossless, comparing random windows of the # converted file against the download python3 - <<'PYEOF' import numpy as np, pyBigWig, random src = pyBigWig.open("/hive/data/outside/proCapNet/hg38/pred/K562.pos.bigWig") out = pyBigWig.open("pred/K562.proCapNet.pos.bw") random.seed(7) bad = 0 for chrom in ("chr1", "chr7", "chr14", "chr21", "chrX"): for _ in range(3): s = random.randrange(0, src.chroms()[chrom] - 200000) a = src.values(chrom, s, s + 200000, numpy=True) b = out.values(chrom, s, s + 200000, numpy=True) na, nb = np.isnan(a), np.isnan(b) if not (na == nb).all() or not (a[~na] == b[~nb]).all(): bad += 1 print("mismatched windows:", bad) PYEOF # mismatched windows: 0 ############################################################################## # negating the minus strand, after the fact ############################################################################## # The prediction files were first built with the minus strand positive and # flipped at display time with the trackDb negateValues setting. That broke the # composite's negate control, as described above, so the twelve minus-strand # files were rewritten with negated values rather than rebuilt from the # downloads, which had already been deleted: # # for db in hg38 hs1 ; do # base=/hive/data/genomes/$db/bed/proCapNet # mkdir -p $base/previousPred # ls $base/pred/*.neg.bw | while read f ; do # ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredToFixedStep \ # --negate /hive/data/genomes/$db/chrom.sizes $f $base/negated/$(basename $f) # done # done # # Each output was checked against its input: same nBasesCovered, value range # mirrored, sampled values exactly negated. The originals were then moved to # previousPred/ and the negated files put in their place, so /gbdb needs no new # symlinks. previousPred/ has since been deleted, 367 GB over the two # assemblies. Nothing is lost by that: negation is its own inverse, so the # originals can be recovered from the files in pred/ by running the same command # again. A rebuild from the downloads does not need any of this either: # proCapNetPredBuild negates the minus strand itself. ############################################################################## # ENCODE 4 PRO-cap ############################################################################## mkdir -p /hive/data/genomes/hg38/bed/encode4ProCap cd /hive/data/genomes/hg38/bed/encode4ProCap # Resolve the portal metadata: for each of the six experiments take only the # plus and minus strand signal of unique reads files belonging to that # experiment's default analysis, which drops files superseded by a later # reprocessing. ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetEncodeMeta \ ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv encodeFiles.tsv # A673 ENCSR046BCI 4 files # Caco-2 ENCSR100LIJ 4 files # Calu3 ENCSR935RNW 4 files # HUVEC ENCSR098LLB 2 files # K562 ENCSR261KBX 4 files # MCF10A ENCSR799DGV 8 files # A copy of the resulting file is committed as # outside/proCapNet/proCapNetEncodeFiles.tsv. It is the record of which ENCODE # file accessions went into each track: re-running the step later can resolve to # different files, because the default analysis of an experiment changes when # ENCODE reprocesses it. # ENCODE publishes no pooled PRO-cap file, so download the per-replicate files # and sum them per cell line and strand. MCF10A has two sequencing runs per # biological replicate, hence eight files. Minus strand signal is negative as # released and is kept that way. ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetEncodeBuild \ encodeFiles.tsv /hive/data/genomes/hg38/chrom.sizes \ /hive/data/outside/proCapNet/encode/raw . # K562.ENCSR261KBX.pos.bw: 2 inputs, 3227685 input intervals, 3051479 merged intervals, # 3194212 bases covered # every input interval is accounted for; total signal is conserved exactly # sanity check: the summed signal equals the sum of the inputs python3 - <<'PYEOF' import pyBigWig def total(f): bw = pyBigWig.open(f); s = bw.header()["sumData"]; bw.close(); return s raw = "/hive/data/outside/proCapNet/encode/raw" print(total(raw + "/ENCFF994CSC.pos.bigWig") + total(raw + "/ENCFF580QEK.pos.bigWig"), total("K562.ENCSR261KBX.pos.bw")) PYEOF # 9605200.0 9605200.0 ############################################################################## # trackDb ############################################################################## # The stanzas are generated. These are traditional composites: a faceted one is # routed to facetedCompositeUi(), which returns before the wiggle controls are # drawn, and signal tracks need those. The matrix is declared over the leaf # bigWigs, so each strand is a cell; the multiWig containers carry no subGroups # because a container spans both strands. cd ~/kent/src/hg/makeDb/outside/proCapNet ./proCapNetTrackDb hg38 proCapNetExperiments.tsv \ ~/kent/src/hg/makeDb/trackDb/human/hg38/transcriptionStart.ra # transcriptionStart.ra is generated; the include line in human/hg38/trackDb.ra # is added by hand: # include transcriptionStart.ra alpha ############################################################################## # cleanup ############################################################################## # The downloads are only needed to build the tracks: the predictions for the # re-encoding, the ENCODE per-replicate files for the summing. /hive charges # twice the apparent size, so this returns about 580 GB of disk. rm -rf /hive/data/outside/proCapNet/hg38/pred /hive/data/outside/proCapNet/encode