acf20930c5cdfa1e352826a702ad8f36344178c4
markd
  Thu Sep 24 21:16:13 2026 -0700
Fixes from an independent review of the TSS tracks. refs #35528

The Data Access sections told users to pass the composite name to the API, which
returns HTTP 400. The API serves one bigWig at a time, so both pages now name a
single strand of one cell line, verified to return 200.

encode4ProCap.html claimed the kent tree held the manifest of which ENCODE files
went into each track, and it did not. Commit that manifest as
proCapNetEncodeFiles.tsv and name it on the page. It matters because
proCapNetEncodeMeta resolves each experiment's default analysis at run time, so
re-running it after an ENCODE reprocessing can pick different files.

Cite Shah et al. for the ENCODE 4 nascent transcriptome survey the six PRO-cap
experiments come from. Sagar Shah was credited by name with no reference.

The hg38 makedoc called the 164268582 dropped bases "the N regions", but gap on
the primary chromosomes is 150610728. The extra 13.7 Mb is sequence flanking each
gap, dropped because most of its 2114 bp window was unresolved.

diff --git src/hg/makeDb/trackDb/human/encode4ProCap.html src/hg/makeDb/trackDb/human/encode4ProCap.html
index 4456dffb38e..41a57dc0988 100644
--- src/hg/makeDb/trackDb/human/encode4ProCap.html
+++ src/hg/makeDb/trackDb/human/encode4ProCap.html
@@ -42,63 +42,68 @@
 </ul>
 
 <h2>Methods</h2>
 
 <p>
 PRO-cap, in the CoPRO form described by Tome <em>et al</em>., permeabilizes cells
 and runs a nuclear run-on reaction with biotinylated NTPs, which RNA polymerase
 incorporates into the nascent transcripts it is actively making. Biotinylated
 RNAs are captured on streptavidin beads, selected for a 5' cap so that only
 transcripts still carrying their original start are kept, then reverse
 transcribed and paired-end sequenced. The 5' end of the second read marks the
 base at which transcription started, and the read's strand gives the strand of
 initiation. The six experiments here were produced by the Yu and Lis labs at
 Cornell University and processed through the ENCODE PRO-cap pipeline, which maps
 reads, removes PCR duplicates using unique molecular identifiers, and writes
-single-base plus and minus strand signal for the uniquely mapping reads.
+single-base plus and minus strand signal for the uniquely mapping reads. These
+experiments are part of the ENCODE 4 nascent transcriptome survey described in
+Shah <em>et al</em>.
 </p>
 
 <p>
 The signal files were downloaded from the
 <a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a>, taking
 for each experiment only the plus and minus strand signal of unique reads files
 belonging to that experiment's default analysis, which drops files superseded by
 a later reprocessing. The experiments are ENCSR046BCI, ENCSR100LIJ, ENCSR935RNW,
 ENCSR098LLB, ENCSR261KBX and ENCSR799DGV for A673, Caco-2, Calu3, HUVEC, K562 and
 MCF10A respectively, and are linked from the Experiment column of the table on
 this page. ENCODE releases PRO-cap signal per replicate and, for MCF10A, per
 sequencing run, and publishes no pooled file, so the files were summed at UCSC to
 give one track per cell line per strand: two files for A673, Caco-2, Calu3 and
 K562, one for HUVEC, which has a single released replicate, and four for MCF10A.
 This is the same merge the ProCapNet models were trained on. Minus strand signal
 is negative as released by ENCODE and is kept that way. Total signal is
 conserved exactly by the summing. The steps are recorded in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/transcriptionStart.txt"
 target="_blank">makeDoc</a> and the scripts are in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/outside/proCapNet"
-target="_blank">the kent source tree</a>, including the manifest recording exactly
-which ENCODE files went into each track.
+target="_blank">the kent source tree</a>, including
+<tt>proCapNetEncodeFiles.tsv</tt>, which records exactly which ENCODE file
+accessions went into each track.
 </p>
 
 <h2>Data Access</h2>
 
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to
 spreadsheet or tab-sep tables. From scripts, the data can be accessed through our
-<a href="https://api.genome.ucsc.edu" target="_blank">API</a>, track=<i>encode4ProCap</i>.
+<a href="https://api.genome.ucsc.edu" target="_blank">API</a>. The API returns one
+bigWig at a time, so name a single strand of one cell line rather than the
+container, for example track=<i>encode4ProCap_K562_procap_pos</i>.
 </p>
 
 <p>
 The Files column of the table on this page links each cell line's bigWigs
 directly, so a single file can be fetched without working out its path.
 </p>
 
 <p>
 For automated download and analysis, the genome annotation is stored in bigWig
 files that can be downloaded from
 <a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4ProCap/" target="_blank">our
 download server</a>. The files are named for the cell line, the ENCODE experiment accession and the
 strand, for example <tt>K562.ENCSR261KBX.pos.bw</tt> and
 <tt>K562.ENCSR261KBX.neg.bw</tt>. Individual regions or the whole
 genome annotation can be obtained using our tool <tt>bigWigToBedGraph</tt>, which
@@ -115,30 +120,44 @@
 <a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a> under the
 experiment accessions listed above.
 </p>
 
 <h2>Credits</h2>
 
 <p>
 PRO-cap data was generated as part of ENCODE by Sagar Shah and the Yu and Lis labs
 at Cornell University. Thanks to the ENCODE Consortium and the ENCODE production
 laboratories for making the data available.
 </p>
 
 <h2>References</h2>
 
 
+<p>
+Shah SR, Chen Y, Leung AK, Navarro PVC, Paramo MI, Gupta J, Gurumurthy A, Fite RF, Weimer AK,
+Cochran K <em>et al</em>.
+<a href="https://www.ncbi.nlm.nih.gov/pubmed/41040139" target="_blank">
+The nascent transcriptome delineates the regulatory landscape in human health and disease</a>.
+<em>bioRxiv</em>. 2026 Jun 9;.
+DOI: <a href="https://doi.org/10.1101/2025.09.24.676871"
+target="_blank">10.1101/2025.09.24.676871</a>; PMID: <a
+href="https://www.ncbi.nlm.nih.gov/pubmed/41040139" target="_blank">41040139</a>; PMC: <a
+href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12485982/" target="_blank">PMC12485982</a>
+</p>
+
+
+
 <p>
 Tome JM, Tippens ND, Lis JT.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/30349116" target="_blank">
 Single-molecule nascent RNA sequencing identifies regulatory domain architecture at promoters and
 enhancers</a>.
 <em>Nat Genet</em>. 2018 Nov;50(11):1533-1541.
 DOI: <a href="https://doi.org/10.1038/s41588-018-0234-5"
 target="_blank">10.1038/s41588-018-0234-5</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/30349116" target="_blank">30349116</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6422046/" target="_blank">PMC6422046</a>
 </p>
 
 
 
 <p>