acf20930c5cdfa1e352826a702ad8f36344178c4
markd
  Thu Sep 24 21:16:13 2026 -0700
Fixes from an independent review of the TSS tracks. refs #35528

The Data Access sections told users to pass the composite name to the API, which
returns HTTP 400. The API serves one bigWig at a time, so both pages now name a
single strand of one cell line, verified to return 200.

encode4ProCap.html claimed the kent tree held the manifest of which ENCODE files
went into each track, and it did not. Commit that manifest as
proCapNetEncodeFiles.tsv and name it on the page. It matters because
proCapNetEncodeMeta resolves each experiment's default analysis at run time, so
re-running it after an ENCODE reprocessing can pick different files.

Cite Shah et al. for the ENCODE 4 nascent transcriptome survey the six PRO-cap
experiments come from. Sagar Shah was credited by name with no reference.

The hg38 makedoc called the 164268582 dropped bases "the N regions", but gap on
the primary chromosomes is 150610728. The extra 13.7 Mb is sequence flanking each
gap, dropped because most of its 2114 bp window was unresolved.

diff --git src/hg/makeDb/doc/hg38/transcriptionStart.txt src/hg/makeDb/doc/hg38/transcriptionStart.txt
index c56e9845456..daa0b13c256 100644
--- src/hg/makeDb/doc/hg38/transcriptionStart.txt
+++ src/hg/makeDb/doc/hg38/transcriptionStart.txt
@@ -15,31 +15,34 @@
 mkdir -p /hive/data/outside/proCapNet/hg38 /hive/data/genomes/hg38/bed/proCapNet
 cd /hive/data/genomes/hg38/bed/proCapNet
 ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetDownload hg38 \
     ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv \
     /hive/data/outside/proCapNet/hg38/pred contrib
 # 290 GB, about 25 minutes at six parallel streams
 
 # Re-encode the predictions from one bedGraph interval per base into fixedStep
 # sections, which cuts the size by about a third without changing a value, and
 # drop the bases holding a literal NaN.  Those NaNs are all in unresolved (N)
 # reference sequence and make the track's autoScale and summary statistics NaN.
 ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredBuild \
     /hive/data/genomes/hg38/chrom.sizes /hive/data/outside/proCapNet/hg38/pred pred 6
 # all 12 files report the same counts:
 #   24 chroms, 3088269832 bases, 2924001250 with data,
-#   164268582 dropped as no-data or NaN   (5.3%, the N regions)
+#   164268582 dropped as no-data or NaN   (5.3%)
+# That is more than the 150610728 gap bases on the primary chromosomes, because
+# no prediction was made where most of a 2114 bp window was unresolved, which
+# also takes out sequence flanking each gap.
 # about 24 GB in, 15.7 to 15.9 GB out per file
 
 # confirm the summary statistics are numbers again, they were nan before
 bigWigInfo pred/K562.proCapNet.pos.bw | egrep 'basesCovered|mean|std'
 #   basesCovered: 2,924,001,250
 #   mean: 0.019478
 #   std: 0.300502
 
 # spot check that the re-encoding is lossless, comparing random windows of the
 # converted file against the download
 python3 - <<'PYEOF'
 import numpy as np, pyBigWig, random
 src = pyBigWig.open("/hive/data/outside/proCapNet/hg38/pred/K562.pos.bigWig")
 out = pyBigWig.open("pred/K562.proCapNet.pos.bw")
 random.seed(7)
@@ -63,30 +66,35 @@
 mkdir -p /hive/data/genomes/hg38/bed/encode4ProCap
 cd /hive/data/genomes/hg38/bed/encode4ProCap
 
 # Resolve the portal metadata: for each of the six experiments take only the
 # plus and minus strand signal of unique reads files belonging to that
 # experiment's default analysis, which drops files superseded by a later
 # reprocessing.
 ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetEncodeMeta \
     ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv encodeFiles.tsv
 #   A673    ENCSR046BCI  4 files
 #   Caco-2  ENCSR100LIJ  4 files
 #   Calu3   ENCSR935RNW  4 files
 #   HUVEC   ENCSR098LLB  2 files
 #   K562    ENCSR261KBX  4 files
 #   MCF10A  ENCSR799DGV  8 files
+# A copy of the resulting file is committed as
+# outside/proCapNet/proCapNetEncodeFiles.tsv.  It is the record of which ENCODE
+# file accessions went into each track: re-running the step later can resolve to
+# different files, because the default analysis of an experiment changes when
+# ENCODE reprocesses it.
 
 # ENCODE publishes no pooled PRO-cap file, so download the per-replicate files
 # and sum them per cell line and strand.  MCF10A has two sequencing runs per
 # biological replicate, hence eight files.  Minus strand signal is negative as
 # released and is kept that way.
 ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetEncodeBuild \
     encodeFiles.tsv /hive/data/genomes/hg38/chrom.sizes \
     /hive/data/outside/proCapNet/encode/raw .
 #   K562.ENCSR261KBX.pos.bw: 2 inputs, 3227685 input intervals, 3051479 merged intervals,
 #                3194212 bases covered
 # every input interval is accounted for; total signal is conserved exactly
 
 # sanity check: the summed signal equals the sum of the inputs
 python3 - <<'PYEOF'
 import pyBigWig