9f6962c173c405e481da17794d67107e7f8e73a9
jnavarr5
  Fri Oct 2 16:05:31 2026 -0700
Replacing the subtrack matrix and Sample class filter paragraph on the PRO-cap and ProCapNet description pages with the track collection layout from Mark's change, and noting that ProCapNet minus strand files store negative values, refs #35528

diff --git src/hg/makeDb/trackDb/human/proCapNet.html src/hg/makeDb/trackDb/human/proCapNet.html
index 463d03e2645..e904555336b 100644
--- src/hg/makeDb/trackDb/human/proCapNet.html
+++ src/hg/makeDb/trackDb/human/proCapNet.html
@@ -1,218 +1,220 @@
 <h2>Description</h2>
 
 <p>
 ProCapNet is a neural network trained to predict PRO-cap signal from DNA
 sequence alone. PRO-cap is a run-on assay that captures the 5' end of each
 nascent RNA, so it reports the exact base and strand at which RNA polymerase II
 started transcribing, including at enhancers and at unstable transcripts that
 RNA-seq and CAGE miss. Six separate models were trained, one on each of six cell
 lines with ENCODE PRO-cap data.
 </p>
 
 <p>
 This track holds two kinds of output from those models:
 </p>
 
 <ul>
 <li><b>Predicted PRO-cap</b>: what the model expects the PRO-cap signal to be,
 at every base of the primary chromosomes, on both strands. The sequence rules that govern
 where initiation happens are largely shared between cell types, so any one model
 highlights sequence capable of driving initiation, including at regions where no
 PRO-cap experiment has been done.</li>
 <li><b>Sequence contribution scores</b>: how much each individual base pushed
 the model's prediction up or down. Bases inside a functional element such as a
 TATA box or an initiator carry high scores, and the pattern of high-scoring
 bases often spells out the recognition sequence of a promoter-associated
 transcription factor. These are computed only around MANE Select transcription
 start sites, so they cover about 1% of the genome and the track is empty
 everywhere else.</li>
 </ul>
 
 <p>
 The predictions are available on GRCh38/hg38 and T2T-CHM13/hs1. The contribution
 scores are available on GRCh38/hg38 only, because they are computed at MANE
 Select transcription start sites and MANE is not defined for T2T-CHM13.
 </p>
 
 <p>
 Predictions are not measurements: they say what the sequence looks capable of,
 not what a given cell is doing. The matching experimental data is in the PRO-cap
 track, available on GRCh38/hg38.
 </p>
 
 <h2>Display Conventions and Configuration</h2>
 
 <p>
-The matrix on this page has one row per cell line and one column per data type,
-so a checkbox turns on one strand of one cell line's predictions, or its
-contribution scores. Use the <b>Sample class</b> filter to restrict the matrix to
-cancer or non-cancer lines.
+This track collection holds one predicted PRO-cap track per cell line, shown by
+default, and on GRCh38/hg38 one contribution score track per cell line, hidden by
+default. Use the buttons on this page to turn tracks on or off, and click a track's
+name to open its own settings, where the plus and minus strands of a predicted
+PRO-cap track can also be turned on separately.
 </p>
 
 <p>
 Each predicted PRO-cap track is an overlay of the two strands: plus strand
 predictions are drawn upward and minus strand predictions downward. The y axis
 is the predicted number of PRO-cap reads at that base.
 </p>
 
 <p>
 Contribution scores are drawn as a sequence logo when zoomed in far enough to
 show individual bases: the letter of the reference base is scaled by its score,
 so a run of tall letters is a motif the model relied on. At lower zoom the same
 values are drawn as a wiggle. Scores can be negative, meaning the base argued
 against initiation being placed where it was. Scores exist only in the roughly
 2 kb window around each MANE Select transcription start site, about 38.7 Mb of
 the genome. Everywhere else the track is empty, which is not the same as a score
 of zero.
 </p>
 
 <p>
 Read depth differs between the six PRO-cap experiments the models were trained
 on, and both predicted signal and contribution scores scale with it, so the
 y axis is not comparable between cell lines. Tracks are colored by the cell line
 the model was trained on:
 </p>
 <ul>
 <li><span style="display:inline-block; background-color:#0072B2; width:18px; height:12px; vertical-align:middle;"></span> <b>A673</b> Ewing sarcoma</li>
 <li><span style="display:inline-block; background-color:#D55E00; width:18px; height:12px; vertical-align:middle;"></span> <b>Caco-2</b> colorectal adenocarcinoma</li>
 <li><span style="display:inline-block; background-color:#009E73; width:18px; height:12px; vertical-align:middle;"></span> <b>Calu3</b> lung adenocarcinoma</li>
 <li><span style="display:inline-block; background-color:#CC79A7; width:18px; height:12px; vertical-align:middle;"></span> <b>HUVEC</b> umbilical vein endothelial cells</li>
 <li><span style="display:inline-block; background-color:#E69F00; width:18px; height:12px; vertical-align:middle;"></span> <b>K562</b> chronic myelogenous leukemia</li>
 <li><span style="display:inline-block; background-color:#56B4E9; width:18px; height:12px; vertical-align:middle;"></span> <b>MCF10A</b> non-tumorigenic breast epithelium</li>
 </ul>
 
 <p>
 One model was trained per cell line. Each links to its ENCODE annotation,
 which holds the trained model itself.
 </p>
 
 <table class="stdTbl">
 <tr><th>Cell line</th><th>Tissue</th><th>Sample class</th><th>ENCODE model</th></tr>
 <tr><td>A673</td><td>Muscle</td><td>Cancer</td><td><a href="https://www.encodeproject.org/annotations/ENCSR072YCM/" target="_blank">ENCSR072YCM</a></td></tr>
 <tr><td>Caco-2</td><td>Colon</td><td>Cancer</td><td><a href="https://www.encodeproject.org/annotations/ENCSR182QNJ/" target="_blank">ENCSR182QNJ</a></td></tr>
 <tr><td>Calu3</td><td>Lung</td><td>Cancer</td><td><a href="https://www.encodeproject.org/annotations/ENCSR797DEF/" target="_blank">ENCSR797DEF</a></td></tr>
 <tr><td>HUVEC</td><td>Blood vessel</td><td>Non-cancer</td><td><a href="https://www.encodeproject.org/annotations/ENCSR801ECP/" target="_blank">ENCSR801ECP</a></td></tr>
 <tr><td>K562</td><td>Blood</td><td>Cancer</td><td><a href="https://www.encodeproject.org/annotations/ENCSR740IPL/" target="_blank">ENCSR740IPL</a></td></tr>
 <tr><td>MCF10A</td><td>Breast</td><td>Non-cancer</td><td><a href="https://www.encodeproject.org/annotations/ENCSR860TYZ/" target="_blank">ENCSR860TYZ</a></td></tr>
 </table>
 
 <h2>Methods</h2>
 
 <p>
 ProCapNet adapts the BPNet architecture: it reads 2,114 bp of sequence and
 predicts a base-resolution initiation profile over the central 1,000 bp on both
 strands, together with the total number of initiation events in that window. One
 model was trained per cell line on ENCODE PRO-cap data, using all PRO-cap peaks
 in that cell line plus sampled DNase-hypersensitive sites as background, with
 7-fold cross-validation split by chromosome. Full details are in Cochran
 <em>et al.</em>, 2024.
 </p>
 
 <p>
 The Kundaje lab generated the genome-wide predictions by applying each model in
 2,114 bp windows at a stride of 250 bp, then averaging across the seven
 cross-validation models and across the forward and reverse-complemented
 sequence. On hg38 no prediction was made where most of a window was unresolved
 (N) in the reference. Contribution scores were computed with DeepSHAP, which
 contrasts the model's output on the real sequence against its output on 25
 dinucleotide shuffles of it, and were computed only at MANE Select transcription
 start sites.
 </p>
 
 <p>
 The ProCapNet model implementation is at
 <a href="https://github.com/kundajelab/ProCapNet" target="_blank">kundajelab/ProCapNet</a>.
 The steps that turned the published files into these tracks are recorded in
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/$db/transcriptionStart.txt"
 target="_blank">doc/$db/transcriptionStart.txt</a>, the scripts they run are in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/outside/proCapNet"
 target="_blank">makeDb/outside/proCapNet</a>, and the track configuration is in
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/$db/transcriptionStart.ra"
 target="_blank">trackDb/human/$db/transcriptionStart.ra</a>.
 </p>
 
 <h2>Data Access</h2>
 
 <p>
 The bigWig files are on our
 <a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/proCapNet/" target="_blank">download
 server</a>. Predictions are under <tt>pred/</tt> and are named for the cell line,
 the model and the strand, for example <tt>K562.proCapNet.pos.bw</tt> and
-<tt>K562.proCapNet.neg.bw</tt>. Contribution scores, which exist for GRCh38 only,
+<tt>K562.proCapNet.neg.bw</tt>. Minus strand values are stored as negative
+numbers. Contribution scores, which exist for GRCh38 only,
 are under <tt>contrib/</tt>, for example <tt>K562.proCapNet-contrib.bw</tt>.
 </p>
 
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to
 spreadsheet or tab-sep tables. From scripts, the data can be accessed through our
 <a href="https://api.genome.ucsc.edu" target="_blank">API</a>. The API returns one
 bigWig at a time, so name a single strand of one cell line rather than the
 container, for example track=<i>proCapNet_K562_pred_pos</i>.
 </p>
 
 <p>
 Individual regions or the whole genome annotation can be obtained using our tool
 <tt>bigWigToBedGraph</tt>, which can be compiled from the source code or
 downloaded as a precompiled binary for your system. Instructions for downloading
 source code and binaries are on the
 <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads"
 target="_blank">utilities download page</a>. The tool can also be used to obtain
 features within a given range, e.g.
 <tt>bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/$db/proCapNet/pred/K562.proCapNet.pos.bw
 -chrom=chr21 -start=0 -end=100000000 stdout</tt>
 </p>
 
 <p>
 The ProCapNet models are on the
 <a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a> as
 BPNet-model annotations, one per cell line, linked from the <b>ENCODE model</b>
 column of the table above. Each annotation also holds the trained model, sequence
 contribution scores and predicted signal over a selected set of regions. The
 genome-wide predictions shown here are not part of that ENCODE release.
 </p>
 
 <h2>Credits</h2>
 
 <p>
 ProCapNet was developed by Kelly Cochran in the Kundaje lab at Stanford
 University. The genome-wide predictions and contribution scores were generated by
 Kelly Cochran in collaboration with the GENCODE consortium. Thanks to Kelly
 Cochran and Anshul Kundaje for making the data available.
 </p>
 
 <h2>References</h2>
 
 <p>
 Cochran K, Yin M, Mantripragada A, Schreiber J, Marinov GK, Shah SR, Yu H, Lis JT, Kundaje A.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/38853896" target="_blank">
 Dissecting the cis-regulatory syntax of transcription initiation with deep learning</a>.
 <em>bioRxiv</em>. 2024 Nov 21;.
 DOI: <a href="https://doi.org/10.1101/2024.05.28.596138"
 target="_blank">10.1101/2024.05.28.596138</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/38853896" target="_blank">38853896</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11160661/" target="_blank">PMC11160661</a>
 </p>
 
 <p>
 Avsec &#381;, Weilert M, Shrikumar A, Krueger S, Alexandari A, Dalal K, Fropf R, McAnany C, Gagneur J,
 Kundaje A <em>et al</em>.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/33603233" target="_blank">
 Base-resolution models of transcription-factor binding reveal soft motif syntax</a>.
 <em>Nat Genet</em>. 2021 Mar;53(3):354-366.
 DOI: <a href="https://doi.org/10.1038/s41588-021-00782-6"
 target="_blank">10.1038/s41588-021-00782-6</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/33603233" target="_blank">33603233</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8812996/" target="_blank">PMC8812996</a>
 </p>
 
 <p>
 Kwak H, Fuda NJ, Core LJ, Lis JT.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/23430654" target="_blank">
 Precise maps of RNA polymerase reveal how promoters direct initiation and pausing</a>.
 <em>Science</em>. 2013 Feb 22;339(6122):950-3.
 DOI: <a href="https://doi.org/10.1126/science.1229386" target="_blank">10.1126/science.1229386</a>;
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/23430654" target="_blank">23430654</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3974810/" target="_blank">PMC3974810</a>
 </p>