fe4c75c272ca27de55c0b42396c47f2b00637a4a
jnavarr5
  Mon Sep 28 16:36:20 2026 -0700
Linking the encode4ProCap trackDb.ra, saying why PRO-cap and contribution scores are hg38 only on the container page, and limiting the ProCapNet prediction claim to the primary chromosomes, refs #35528

diff --git src/hg/makeDb/trackDb/human/encode4ProCap.html src/hg/makeDb/trackDb/human/encode4ProCap.html
index 1a5f1b8a93c..e6831ba0819 100644
--- src/hg/makeDb/trackDb/human/encode4ProCap.html
+++ src/hg/makeDb/trackDb/human/encode4ProCap.html
@@ -1,169 +1,171 @@
 <h2>Description</h2>
 
 <p>
 Transcription begins when RNA polymerase II starts making an RNA at a
 transcription start site. A promoter or an enhancer does not use a single start
 site: polymerase initiates across a spread of nearby bases, on both strands, and
 which bases are used is set largely by the local DNA sequence. PRO-cap is a
 nascent RNA run-on assay that captures the 5' end of each new transcript, so each
 read reports the exact base and strand of one initiation event. Unlike RNA-seq or
 CAGE, PRO-cap sees unstable RNAs as well as stable ones, which makes initiation
 at enhancers visible.
 </p>
 
 <p>
 This track shows PRO-cap signal from six ENCODE 4 experiments, one per cell line,
 on GRCh38/hg38 only. The predictions a deep learning model makes from sequence
 alone, trained on this same data, are in the ProCapNet track, available on
 GRCh38/hg38 and T2T-CHM13/hs1.
 </p>
 
 <h2>Display Conventions and Configuration</h2>
 
 <p>
 Cell lines are listed in the table on this page, one row each. Use the <b>Sample
 class</b> facet to narrow the list, and the <b>Experiment</b> column to open the experiment
 on the ENCODE portal.
 </p>
 
 <p>
 Each track is an overlay of the two strands: plus strand reads are drawn upward
 and minus strand reads downward. The y axis is the number of initiation events
 measured at that base, summed over the experiment's replicates, so tracks with
 deeper sequencing reach higher values. Tracks are colored by cell line:
 </p>
 <ul>
 <li><span style="display:inline-block; background-color:#0072B2; width:18px; height:12px; vertical-align:middle;"></span> <b>A673</b> Ewing sarcoma</li>
 <li><span style="display:inline-block; background-color:#D55E00; width:18px; height:12px; vertical-align:middle;"></span> <b>Caco-2</b> colorectal adenocarcinoma</li>
 <li><span style="display:inline-block; background-color:#009E73; width:18px; height:12px; vertical-align:middle;"></span> <b>Calu3</b> lung adenocarcinoma</li>
 <li><span style="display:inline-block; background-color:#CC79A7; width:18px; height:12px; vertical-align:middle;"></span> <b>HUVEC</b> umbilical vein endothelial cells</li>
 <li><span style="display:inline-block; background-color:#E69F00; width:18px; height:12px; vertical-align:middle;"></span> <b>K562</b> chronic myelogenous leukemia</li>
 <li><span style="display:inline-block; background-color:#56B4E9; width:18px; height:12px; vertical-align:middle;"></span> <b>MCF10A</b> non-tumorigenic breast epithelium</li>
 </ul>
 
 <h2>Methods</h2>
 
 <p>
 PRO-cap, in the CoPRO form of Tome <em>et al.</em>, 2018, runs a nuclear run-on
 reaction with biotinylated NTPs, captures the biotinylated nascent RNAs, selects
 for a 5' cap so that only transcripts carrying their original start are kept, and
 sequences them paired-end. The 5' end of the second read marks the base at which
 transcription started and its strand gives the strand of initiation. The six
 experiments were produced by the Yu and Lis labs at Cornell University as part of
 the ENCODE 4 nascent transcriptome survey described in Shah <em>et al.</em>,
 2026, and processed through the ENCODE PRO-cap pipeline.
 </p>
 
 <p>
 The signal files were downloaded from the
 <a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a>, taking
 only the plus and minus strand signal of unique reads belonging to each
 experiment's default analysis. ENCODE publishes no pooled file, so the
 per-replicate files were summed at UCSC into one track per cell line and strand,
 the same merge the ProCapNet models were trained on. Total signal is conserved
 exactly by the summing, and minus strand signal is negative as released. The
 experiments are ENCSR046BCI, ENCSR100LIJ, ENCSR935RNW, ENCSR098LLB, ENCSR261KBX
 and ENCSR799DGV for A673, Caco-2, Calu3, HUVEC, K562 and MCF10A respectively.
 </p>
 
 <p>
 The steps are recorded in
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/transcriptionStart.txt"
-target="_blank">doc/hg38/transcriptionStart.txt</a> and the scripts are in
+target="_blank">doc/hg38/transcriptionStart.txt</a>, the scripts are in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/outside/proCapNet"
 target="_blank">makeDb/outside/proCapNet</a>, including
 <tt>proCapNetEncodeFiles.tsv</tt>, which records exactly which ENCODE file
-accessions went into each track.
+accessions went into each track, and the track configuration is in
+<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/hg38/transcriptionStart.ra"
+target="_blank">trackDb/human/hg38/transcriptionStart.ra</a>.
 </p>
 
 <h2>Data Access</h2>
 
 <p>
 The bigWig files are on our
 <a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4ProCap/" target="_blank">download
 server</a>, named for the cell line, the ENCODE experiment accession and the
 strand, for example <tt>K562.ENCSR261KBX.pos.bw</tt> and
 <tt>K562.ENCSR261KBX.neg.bw</tt>.
 </p>
 
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to
 spreadsheet or tab-sep tables. From scripts, the data can be accessed through our
 <a href="https://api.genome.ucsc.edu" target="_blank">API</a>. The API returns one
 bigWig at a time, so name a single strand of one cell line rather than the
 container, for example track=<i>encode4ProCap_K562_procap_pos</i>.
 </p>
 
 <p>
 Individual regions or the whole genome annotation can be obtained using our tool
 <tt>bigWigToBedGraph</tt>, which can be compiled from the source code or
 downloaded as a precompiled binary for your system. Instructions for downloading
 source code and binaries are on the
 <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads"
 target="_blank">utilities download page</a>. The tool can also be used to obtain
 features within a given range, e.g.
 <tt>bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4ProCap/K562.ENCSR261KBX.pos.bw
 -chrom=chr21 -start=0 -end=100000000 stdout</tt>
 </p>
 
 <p>
 The unmerged per-replicate files can be downloaded from the
 <a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a> under the
 experiment accessions listed above.
 </p>
 
 <h2>Credits</h2>
 
 <p>
 PRO-cap data was generated as part of ENCODE by Sagar Shah and the Yu and Lis labs
 at Cornell University. Thanks to the ENCODE Consortium and the ENCODE production
 laboratories for making the data available.
 </p>
 
 <h2>References</h2>
 
 <p>
 Shah SR, Chen Y, Leung AK, Navarro PVC, Paramo MI, Gupta J, Gurumurthy A, Fite RF, Weimer AK,
 Cochran K <em>et al</em>.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/41040139" target="_blank">
 The nascent transcriptome delineates the regulatory landscape in human health and disease</a>.
 <em>bioRxiv</em>. 2026 Jun 9;.
 DOI: <a href="https://doi.org/10.1101/2025.09.24.676871"
 target="_blank">10.1101/2025.09.24.676871</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/41040139" target="_blank">41040139</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12485982/" target="_blank">PMC12485982</a>
 </p>
 
 <p>
 Tome JM, Tippens ND, Lis JT.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/30349116" target="_blank">
 Single-molecule nascent RNA sequencing identifies regulatory domain architecture at promoters and
 enhancers</a>.
 <em>Nat Genet</em>. 2018 Nov;50(11):1533-1541.
 DOI: <a href="https://doi.org/10.1038/s41588-018-0234-5"
 target="_blank">10.1038/s41588-018-0234-5</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/30349116" target="_blank">30349116</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6422046/" target="_blank">PMC6422046</a>
 </p>
 
 <p>
 Kwak H, Fuda NJ, Core LJ, Lis JT.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/23430654" target="_blank">
 Precise maps of RNA polymerase reveal how promoters direct initiation and pausing</a>.
 <em>Science</em>. 2013 Feb 22;339(6122):950-3.
 DOI: <a href="https://doi.org/10.1126/science.1229386" target="_blank">10.1126/science.1229386</a>;
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/23430654" target="_blank">23430654</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3974810/" target="_blank">PMC3974810</a>
 </p>
 
 <p>
 Luo Y, Hitz BC, Gabdank I, Hilton JA, Kagda MS, Lam B, Myers Z, Sud P, Jou J, Lin K <em>et al</em>.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/31713622" target="_blank">
 New developments on the Encyclopedia of DNA Elements (ENCODE) data portal</a>.
 <em>Nucleic Acids Res</em>. 2020 Jan 8;48(D1):D882-D889.
 DOI: <a href="https://doi.org/10.1093/nar/gkz1062" target="_blank">10.1093/nar/gkz1062</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/31713622" target="_blank">31713622</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7061942/" target="_blank">PMC7061942</a>
 </p>