1e8f4a189b1ad69e1cc4d60177b74ab1fd249da9
markd
  Fri Oct 2 22:22:36 2026 -0700
Record that previousPred has been deleted. refs #35528

367 GB of un-negated minus-strand predictions, kept while the negated files went
through QA. Deleting them loses nothing recoverable: negation is its own
inverse, so an original comes back by running proCapNetPredToFixedStep --negate
over the file now in pred/.

diff --git src/hg/makeDb/doc/hs1/transcriptionStart.txt src/hg/makeDb/doc/hs1/transcriptionStart.txt
index 628bcd0c8c0..9f5b959ee8b 100644
--- src/hg/makeDb/doc/hs1/transcriptionStart.txt
+++ src/hg/makeDb/doc/hs1/transcriptionStart.txt
@@ -1,91 +1,93 @@
 # Transcription start sites: ProCapNet predictions, #35528, Claude
 Thu Sep 17 2026 (Claude/markd)
 
 # Layout: one superTrack per data source, each holding its multiWig overlays
 # directly, the way fantom5 does.  Not a composite: under a composite hgTrackUi
 # lists descendant leaves rather than containers, so an overlay gets no
 # configuration block of its own and hiding both strands of a cell line leaves an
 # empty row where the overlay was (#38441).  As a superTrack member each overlay
 # is a track in its own right, with a full configuration page, and hiding it
 # hides the whole overlay.  The cost is the subtrack matrix and the sample class
 # filter, neither of which a superTrack offers.
 #
 # These two superTracks are top level rather than sitting inside a single
 # transcriptionStart folder, because superTracks do not nest: a superTrack given
 # a parent passes tdbQuery -check -strict and is then silently dropped at load
 # (#38460).  transcriptionStart.html is left in the tree unused against that
 # being fixed.
 
 # hs1 has the ProCapNet predictions only.  There is no PRO-cap experiment and no
 # sequence-contribution score set on this assembly.  See
 # makeDb/doc/hg38/transcriptionStart.txt for the hg38 build, which carries all
 # three.  Scripts are in ~/kent/src/hg/makeDb/outside/proCapNet.
 
 mkdir -p /hive/data/outside/proCapNet/hs1 /hive/data/genomes/hs1/bed/proCapNet
 cd /hive/data/genomes/hs1/bed/proCapNet
 ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetDownload hs1 \
     ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv \
     /hive/data/outside/proCapNet/hs1/pred
 # 300 GB, about 30 minutes at six parallel streams
 
 # Re-encode from one bedGraph interval per base into fixedStep sections.  Unlike
 # hg38 these files hold no NaN, so every base is kept.
 ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredBuild \
     /hive/data/genomes/hs1/chrom.sizes /hive/data/outside/proCapNet/hs1/pred pred 6
 # all 12 files report:
 #   24 chroms, 3117275501 bases, 3117275501 with data, 0 dropped
 # about 25 GB in, 16.7 to 17.0 GB out per file
 
 # proCapNetPredBuild writes the minus strand negated, so it draws below the
 # baseline from its own values.  Doing it in the data rather than with the
 # trackDb negateValues setting is what makes the composite's own negate control
 # work: that control sets one shared value for the composite, which replaced the
 # per-track negateValues and sent both strands the same way, with no way back to
 # the default short of a cart reset.  encode4ProCap never had the problem,
 # because ENCODE already publishes its minus strand negative.
 
 # the mean matches the published file exactly, so nothing was altered
 bigWigInfo pred/K562.proCapNet.pos.bw | egrep 'basesCovered|mean'
 #   basesCovered: 3,117,275,501
 #   mean: 0.019773
 
 ##############################################################################
 # negating the minus strand, after the fact
 ##############################################################################
 
 # The prediction files were first built with the minus strand positive and
 # flipped at display time with the trackDb negateValues setting.  That broke the
 # composite's negate control, as described above, so the twelve minus-strand
 # files were rewritten with negated values rather than rebuilt from the
 # downloads, which had already been deleted:
 #
 #   for db in hg38 hs1 ; do
 #       base=/hive/data/genomes/$db/bed/proCapNet
 #       mkdir -p $base/previousPred
 #       ls $base/pred/*.neg.bw | while read f ; do
 #           ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredToFixedStep \
 #               --negate /hive/data/genomes/$db/chrom.sizes $f $base/negated/$(basename $f)
 #       done
 #   done
 #
 # Each output was checked against its input: same nBasesCovered, value range
 # mirrored, sampled values exactly negated.  The originals were then moved to
 # previousPred/ and the negated files put in their place, so /gbdb needs no new
-# symlinks.  previousPred/ is about 197 GB and can be removed once the track has
-# been through QA.  A rebuild from the downloads does not need any of this:
+# symlinks.  previousPred/ has since been deleted, 367 GB over the two
+# assemblies.  Nothing is lost by that: negation is its own inverse, so the
+# originals can be recovered from the files in pred/ by running the same command
+# again.  A rebuild from the downloads does not need any of this either:
 # proCapNetPredBuild negates the minus strand itself.
 
 cd ~/kent/src/hg/makeDb/outside/proCapNet
 ./proCapNetTrackDb hs1 proCapNetExperiments.tsv \
     ~/kent/src/hg/makeDb/trackDb/human/hs1/transcriptionStart.ra
 
 # transcriptionStart.ra is generated; the include line in human/hs1/trackDb.ra
 # is added by hand:
 #   include transcriptionStart.ra alpha
 
 # hs1 is a curated hub assembly, so the stanzas reach the browser through the
 # hub built under /gbdb/hs1/hubs/$USER by the trackDb make, and are visible only
 # when curatedHubPrefix in the sandbox hg.conf names that directory.
 
 # the downloads are only needed for the re-encoding
 rm -rf /hive/data/outside/proCapNet/hs1/pred