40264d0668926b51da94d2ca038dc7349dfc3405
markd
  Sat Sep 26 07:59:34 2026 -0700
Shorten the Methods sections on the two TSS data pages. refs #35528

Both read like a paper's methods section rather than a track description: 43
lines over four subsections on proCapNet, 35 on encode4ProCap. Cut each to three
paragraphs, what the model or assay is, how the files we serve were produced, and
where the build is recorded. The subsection headings go with it.

Kept every fact a user of the track needs: the window and stride, the MANE Select
limit on the contribution scores, the unresolved-sequence handling on hg38, the
replicate summing, the six ENCODE accessions and the file manifest. Dropped the
detail that belongs to the papers, such as the DeepSHAP scalarization and the
loss weighting.

The container page keeps no Methods section, which is correct for a page that
only points at the two data pages.

diff --git src/hg/makeDb/trackDb/human/encode4ProCap.html src/hg/makeDb/trackDb/human/encode4ProCap.html
index 1ab67b3c140..91a78676c91 100644
--- src/hg/makeDb/trackDb/human/encode4ProCap.html
+++ src/hg/makeDb/trackDb/human/encode4ProCap.html
@@ -32,64 +32,58 @@
 measured at that base, summed over the experiment's replicates, so tracks with
 deeper sequencing reach higher values. Tracks are colored by cell line:
 </p>
 <ul>
 <li><span style="display:inline-block; background-color:#0072B2; width:18px; height:12px; vertical-align:middle;"></span> <b>A673</b> Ewing sarcoma</li>
 <li><span style="display:inline-block; background-color:#D55E00; width:18px; height:12px; vertical-align:middle;"></span> <b>Caco-2</b> colorectal adenocarcinoma</li>
 <li><span style="display:inline-block; background-color:#009E73; width:18px; height:12px; vertical-align:middle;"></span> <b>Calu3</b> lung adenocarcinoma</li>
 <li><span style="display:inline-block; background-color:#CC79A7; width:18px; height:12px; vertical-align:middle;"></span> <b>HUVEC</b> umbilical vein endothelial cells</li>
 <li><span style="display:inline-block; background-color:#E69F00; width:18px; height:12px; vertical-align:middle;"></span> <b>K562</b> chronic myelogenous leukemia</li>
 <li><span style="display:inline-block; background-color:#56B4E9; width:18px; height:12px; vertical-align:middle;"></span> <b>MCF10A</b> non-tumorigenic breast epithelium</li>
 </ul>
 
 <h2>Methods</h2>
 
 <p>
-PRO-cap, in the CoPRO form described by Tome <em>et al.</em>, 2018, permeabilizes cells
-and runs a nuclear run-on reaction with biotinylated NTPs, which RNA polymerase
-incorporates into the nascent transcripts it is actively making. Biotinylated
-RNAs are captured on streptavidin beads, selected for a 5' cap so that only
-transcripts still carrying their original start are kept, then reverse
-transcribed and paired-end sequenced. The 5' end of the second read marks the
-base at which transcription started, and the read's strand gives the strand of
-initiation. The six experiments here were produced by the Yu and Lis labs at
-Cornell University and processed through the ENCODE PRO-cap pipeline, which maps
-reads, removes PCR duplicates using unique molecular identifiers, and writes
-single-base plus and minus strand signal for the uniquely mapping reads. These
-experiments are part of the ENCODE 4 nascent transcriptome survey described in
-Shah <em>et al.</em>, 2026.
+PRO-cap, in the CoPRO form of Tome <em>et al.</em>, 2018, runs a nuclear run-on
+reaction with biotinylated NTPs, captures the biotinylated nascent RNAs, selects
+for a 5' cap so that only transcripts carrying their original start are kept, and
+sequences them paired-end. The 5' end of the second read marks the base at which
+transcription started and its strand gives the strand of initiation. The six
+experiments were produced by the Yu and Lis labs at Cornell University as part of
+the ENCODE 4 nascent transcriptome survey described in Shah <em>et al.</em>,
+2026, and processed through the ENCODE PRO-cap pipeline.
 </p>
 
 <p>
 The signal files were downloaded from the
 <a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a>, taking
-for each experiment only the plus and minus strand signal of unique reads files
-belonging to that experiment's default analysis, which drops files superseded by
-a later reprocessing. The experiments are ENCSR046BCI, ENCSR100LIJ, ENCSR935RNW,
-ENCSR098LLB, ENCSR261KBX and ENCSR799DGV for A673, Caco-2, Calu3, HUVEC, K562 and
-MCF10A respectively, and are linked from the Experiment column of the table on
-this page. ENCODE releases PRO-cap signal per replicate and, for MCF10A, per
-sequencing run, and publishes no pooled file, so the files were summed at UCSC to
-give one track per cell line per strand: two files for A673, Caco-2, Calu3 and
-K562, one for HUVEC, which has a single released replicate, and four for MCF10A.
-This is the same merge the ProCapNet models were trained on. Minus strand signal
-is negative as released by ENCODE and is kept that way. Total signal is
-conserved exactly by the summing. The steps are recorded in the
+only the plus and minus strand signal of unique reads belonging to each
+experiment's default analysis. ENCODE publishes no pooled file, so the
+per-replicate files were summed at UCSC into one track per cell line and strand,
+the same merge the ProCapNet models were trained on. Total signal is conserved
+exactly by the summing, and minus strand signal is negative as released. The
+experiments are ENCSR046BCI, ENCSR100LIJ, ENCSR935RNW, ENCSR098LLB, ENCSR261KBX
+and ENCSR799DGV for A673, Caco-2, Calu3, HUVEC, K562 and MCF10A respectively.
+</p>
+
+<p>
+The steps are recorded in
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/transcriptionStart.txt"
-target="_blank">makeDoc</a> and the scripts are in
+target="_blank">doc/hg38/transcriptionStart.txt</a> and the scripts are in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/outside/proCapNet"
-target="_blank">the kent source tree</a>, including
+target="_blank">makeDb/outside/proCapNet</a>, including
 <tt>proCapNetEncodeFiles.tsv</tt>, which records exactly which ENCODE file
 accessions went into each track.
 </p>
 
 <h2>Data Access</h2>
 
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to
 spreadsheet or tab-sep tables. From scripts, the data can be accessed through our
 <a href="https://api.genome.ucsc.edu" target="_blank">API</a>. The API returns one
 bigWig at a time, so name a single strand of one cell line rather than the
 container, for example track=<i>encode4ProCap_K562_procap_pos</i>.
 </p>