5e83632c91c1c536880da1e364a41467856135bd
markd
  Tue Sep 22 11:02:32 2026 -0700
Rename the TSS container and trim the ProCapNet description. refs #35528

Both labels on the transcriptionStart container are now "Transcription
Initiation (TSS)". Changed in proCapNetTrackDb and the two transcriptionStart.ra
files regenerated from it, since they are generated and carry a do-not-edit
header.

Dropped the Processing at UCSC section from the ProCapNet page. How the files
reached UCSC is not something a browser user needs; the makeDoc already records
it, including the NaN bases dropped from the hg38 predictions.

Replaced the Kundaje lab server link with the ENCODE portal. The six ENCODE
records are BPNet-model annotations holding the trained model, contribution
scores and predicted signal over a selected region set. The genome-wide
predictions in this track are roughly fifty times larger than the ENCODE
predicted-signal files and are not part of that release, so the page says so
rather than naming ENCODE as their source.

Claude-Session: https://claude.ai/code/session_01LAB6jWshLvB7eNXQKWVuW5

diff --git src/hg/makeDb/trackDb/human/proCapNet.html src/hg/makeDb/trackDb/human/proCapNet.html
index c5670d1a0d7..687cf882d11 100644
--- src/hg/makeDb/trackDb/human/proCapNet.html
+++ src/hg/makeDb/trackDb/human/proCapNet.html
@@ -106,51 +106,30 @@
 
 <p>
 Scores were computed with DeepSHAP, which estimates each base's contribution by
 contrasting the model's output on the real sequence against its output on a set
 of reference sequences, here 25 dinucleotide shuffles of the sequence being
 scored. Because DeepSHAP needs a single scalar to explain, the base-resolution
 profile output was summarized by mean-normalizing the pre-softmax logits and
 taking their dot product with the post-softmax profile, which weights each
 base's logit by its predicted probability of being used and sums over the
 1,000 bp output window and both strands. This is the profile or TSS-positioning
 task; ProCapNet can also produce scores for its read-count task, which are not
 shown here. Each scored sequence was run through all seven cross-validation
 models and in both orientations, and the scores averaged.
 </p>
 
-<h3>Processing at UCSC</h3>
-
-<p>
-The bigWig files were downloaded from
-<a href="https://mitra.stanford.edu/kundaje/kcochran/ucsc_track_hubs_public/ProCapNet/"
-target="_blank">the Kundaje lab server</a>. The contribution score files are used
-unaltered. In the prediction files the per-base values were re-encoded from one
-interval per base into fixedStep sections, which cuts the file size by about a
-third and changes no value. Bases carrying a literal NaN in the published hg38
-files, all of them in unresolved (N) reference sequence, were dropped so that the
-track autoscales and summarizes correctly; this removed 164,268,582 of the
-3,088,269,832 bases on the primary hg38 chromosomes, leaving 2,924,001,250 bases
-with a prediction. The hs1 files had no such bases and all 3,117,275,501 were
-kept. Nothing else was altered: every retained value is bit for bit the published
-value. The steps are recorded in the
-<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/transcriptionStart.txt"
-target="_blank">makeDoc</a> and the scripts are in
-<a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/outside/proCapNet"
-target="_blank">the kent source tree</a>.
-</p>
-
 <h2>Data Access</h2>
 
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to
 spreadsheet or tab-sep tables. From scripts, the data can be accessed through our
 <a href="https://api.genome.ucsc.edu" target="_blank">API</a>, track=<i>proCapNet</i>.
 </p>
 
 <p>
 The Files column of the table on this page links each cell line's bigWigs
 directly, so a single file can be fetched without working out its path.
 </p>
 
@@ -163,34 +142,36 @@
 <tt>K562.proCapNet.pos.bw</tt> and <tt>K562.proCapNet.neg.bw</tt>. Contribution
 scores are under <tt>contrib/</tt>, for example
 <tt>K562.proCapNet-contrib.bw</tt>. The T2T-CHM13 predictions are in the same
 layout under <tt>/gbdb/hs1/proCapNet/pred/</tt>; there are no contribution scores
 for that assembly. Individual regions or the whole genome annotation
 can be obtained using our tool <tt>bigWigToBedGraph</tt>, which can be compiled
 from the source code or downloaded as a precompiled binary for your system.
 Instructions for downloading source code and binaries can be found
 <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" target="_blank">here</a>.
 The tool can also be used to obtain features within a given range, e.g.
 <tt>bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/hg38/proCapNet/pred/K562.proCapNet.pos.bw
 -chrom=chr21 -start=0 -end=100000000 stdout</tt>
 </p>
 
 <p>
-The original files can be downloaded from
-<a href="https://mitra.stanford.edu/kundaje/kcochran/ucsc_track_hubs_public/ProCapNet/"
-target="_blank">the Kundaje lab server</a>. The trained models are on the ENCODE
-portal, linked from the Experiment column of the table on this page.
+The ProCapNet models are on the
+<a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a> as
+BPNet-model annotations, one per cell line, linked from the Experiment column of
+the table on this page. Each annotation also holds the trained model, sequence
+contribution scores and predicted signal over a selected set of regions. The
+genome-wide predictions shown here are not part of that ENCODE release.
 </p>
 
 <h2>Credits</h2>
 
 <p>
 ProCapNet was developed by Kelly Cochran in the Kundaje lab at Stanford
 University. The genome-wide predictions and contribution scores were generated by
 Kelly Cochran in collaboration with the GENCODE consortium. Thanks to Kelly
 Cochran and Anshul Kundaje for making the data available.
 </p>
 
 <h2>References</h2>
 
 
 <p>