0f951bfc96241a7362d1512ceecf59a02682a839
markd
  Thu Sep 24 20:40:52 2026 -0700
Description page fixes from a QA pre-pass on the TSS tracks. refs #35528

Encode the non-ASCII character in the Avsec reference on proCapNet.html.
getTrackReferences emits raw UTF-8, which the browser does not transcode.

Use $db rather than a hardcoded hg38 in the proCapNet download-server link and
the bigWigToBedGraph example. One page serves both assemblies, so an hs1 reader
was being pointed at hg38 files.

Add a Source subsection to proCapNet.html linking the makedoc, the build scripts,
the trackDb file and the upstream kundajelab/ProCapNet repository. These links
went missing when the Processing at UCSC section was dropped; the prose stays
dropped.

Drop the cross-links between the two track pages. Track names are prefixed
hub_<id>_ on hs1, which is a curated hub, and the id is machine-specific, so a
bare-name link cannot work there. Each page now states which assemblies the data
is available on instead, naming both GRCh38/hg38 and T2T-CHM13/hs1.

Bold the two UI control names on the proCapNet display conventions section.

diff --git src/hg/makeDb/trackDb/human/proCapNet.html src/hg/makeDb/trackDb/human/proCapNet.html
index 687cf882d11..b61f4e2cb59 100644
--- src/hg/makeDb/trackDb/human/proCapNet.html
+++ src/hg/makeDb/trackDb/human/proCapNet.html
@@ -12,45 +12,51 @@
 <p>
 This track holds two kinds of output from those models:
 </p>
 
 <ul>
 <li><b>Predicted PRO-cap</b>: what the model expects the PRO-cap signal to be,
 at every base of the genome, on both strands. The sequence rules that govern
 where initiation happens are largely shared between cell types, so any one model
 highlights sequence capable of driving initiation, including at regions where no
 PRO-cap experiment has been done.</li>
 <li><b>Sequence contribution scores</b>: how much each individual base pushed
 the model's prediction up or down. Bases inside a functional element such as a
 TATA box or an initiator carry high scores, and the pattern of high-scoring
 bases often spells out the recognition sequence of a promoter-associated
 transcription factor. These are computed only around MANE Select transcription
-start sites, and exist for GRCh38 only.</li>
+start sites.</li>
 </ul>
 
+<p>
+The predictions are available on GRCh38/hg38 and T2T-CHM13/hs1. The contribution
+scores are available on GRCh38/hg38 only, since that is the assembly they were
+computed against.
+</p>
+
 <p>
 Predictions are not measurements: they say what the sequence looks capable of,
-not what a given cell is doing. The matching experimental data is in the
-<a href="hgTrackUi?g=encode4ProCap">PRO-cap</a> track.
+not what a given cell is doing. The matching experimental data is in the PRO-cap
+track, available on GRCh38/hg38.
 </p>
 
 <h2>Display Conventions and Configuration</h2>
 
 <p>
 Cell lines are listed in the table on this page, one row each, with a checkbox
-per data type. Use the Sample class facet to narrow the list, and the Group
-tracks by buttons to order the browser by sample or by data type.
+per data type. Use the <b>Sample class</b> facet to narrow the list, and the
+<b>Group tracks by</b> buttons to order the browser by sample or by data type.
 </p>
 
 <p>
 Each predicted PRO-cap track is an overlay of the two strands: plus strand
 predictions are drawn upward and minus strand predictions downward. The y axis
 is the predicted number of PRO-cap reads at that base.
 </p>
 
 <p>
 Contribution scores are drawn as a sequence logo when zoomed in far enough to
 show individual bases: the letter of the reference base is scaled by its score,
 so a run of tall letters is a motif the model relied on. At lower zoom the same
 values are drawn as a wiggle. Scores can be negative, meaning the base argued
 against initiation being placed where it was. Scores exist only in the roughly
 2 kb window around each MANE Select transcription start site, about 38.7 Mb of
@@ -106,62 +112,74 @@
 
 <p>
 Scores were computed with DeepSHAP, which estimates each base's contribution by
 contrasting the model's output on the real sequence against its output on a set
 of reference sequences, here 25 dinucleotide shuffles of the sequence being
 scored. Because DeepSHAP needs a single scalar to explain, the base-resolution
 profile output was summarized by mean-normalizing the pre-softmax logits and
 taking their dot product with the post-softmax profile, which weights each
 base's logit by its predicted probability of being used and sums over the
 1,000 bp output window and both strands. This is the profile or TSS-positioning
 task; ProCapNet can also produce scores for its read-count task, which are not
 shown here. Each scored sequence was run through all seven cross-validation
 models and in both orientations, and the scores averaged.
 </p>
 
+<h3>Source</h3>
+
+<p>
+The ProCapNet model implementation is at
+<a href="https://github.com/kundajelab/ProCapNet" target="_blank">kundajelab/ProCapNet</a>.
+The steps that turned the published files into these tracks are recorded in
+<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/$db/transcriptionStart.txt"
+target="_blank">doc/$db/transcriptionStart.txt</a>, the scripts they run are in
+<a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/outside/proCapNet"
+target="_blank">makeDb/outside/proCapNet</a>, and the track configuration is in
+<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/$db/transcriptionStart.ra"
+target="_blank">trackDb/human/$db/transcriptionStart.ra</a>.
+</p>
+
 <h2>Data Access</h2>
 
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to
 spreadsheet or tab-sep tables. From scripts, the data can be accessed through our
 <a href="https://api.genome.ucsc.edu" target="_blank">API</a>, track=<i>proCapNet</i>.
 </p>
 
 <p>
 The Files column of the table on this page links each cell line's bigWigs
 directly, so a single file can be fetched without working out its path.
 </p>
 
 <p>
 For automated download and analysis, the genome annotation is stored in bigWig
 files that can be downloaded from
-<a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/proCapNet/" target="_blank">our
+<a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/proCapNet/" target="_blank">our
 download server</a>. Predictions are under <tt>pred/</tt> and are named for the
 cell line, the model and the strand, for example
 <tt>K562.proCapNet.pos.bw</tt> and <tt>K562.proCapNet.neg.bw</tt>. Contribution
-scores are under <tt>contrib/</tt>, for example
-<tt>K562.proCapNet-contrib.bw</tt>. The T2T-CHM13 predictions are in the same
-layout under <tt>/gbdb/hs1/proCapNet/pred/</tt>; there are no contribution scores
-for that assembly. Individual regions or the whole genome annotation
+scores, which exist for GRCh38 only, are under <tt>contrib/</tt>, for example
+<tt>K562.proCapNet-contrib.bw</tt>. Individual regions or the whole genome annotation
 can be obtained using our tool <tt>bigWigToBedGraph</tt>, which can be compiled
 from the source code or downloaded as a precompiled binary for your system.
 Instructions for downloading source code and binaries can be found
 <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" target="_blank">here</a>.
 The tool can also be used to obtain features within a given range, e.g.
-<tt>bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/hg38/proCapNet/pred/K562.proCapNet.pos.bw
+<tt>bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/$db/proCapNet/pred/K562.proCapNet.pos.bw
 -chrom=chr21 -start=0 -end=100000000 stdout</tt>
 </p>
 
 <p>
 The ProCapNet models are on the
 <a href="https://www.encodeproject.org" target="_blank">ENCODE portal</a> as
 BPNet-model annotations, one per cell line, linked from the Experiment column of
 the table on this page. Each annotation also holds the trained model, sequence
 contribution scores and predicted signal over a selected set of regions. The
 genome-wide predictions shown here are not part of that ENCODE release.
 </p>
 
 <h2>Credits</h2>
 
 <p>
@@ -176,31 +194,31 @@
 
 <p>
 Cochran K, Yin M, Mantripragada A, Schreiber J, Marinov GK, Shah SR, Yu H, Lis JT, Kundaje A.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/38853896" target="_blank">
 Dissecting the cis-regulatory syntax of transcription initiation with deep learning</a>.
 <em>bioRxiv</em>. 2024 Nov 21;.
 DOI: <a href="https://doi.org/10.1101/2024.05.28.596138"
 target="_blank">10.1101/2024.05.28.596138</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/38853896" target="_blank">38853896</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11160661/" target="_blank">PMC11160661</a>
 </p>
 
 
 
 <p>
-Avsec Ž, Weilert M, Shrikumar A, Krueger S, Alexandari A, Dalal K, Fropf R, McAnany C, Gagneur J,
+Avsec &#381;, Weilert M, Shrikumar A, Krueger S, Alexandari A, Dalal K, Fropf R, McAnany C, Gagneur J,
 Kundaje A <em>et al</em>.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/33603233" target="_blank">
 Base-resolution models of transcription-factor binding reveal soft motif syntax</a>.
 <em>Nat Genet</em>. 2021 Mar;53(3):354-366.
 DOI: <a href="https://doi.org/10.1038/s41588-021-00782-6"
 target="_blank">10.1038/s41588-021-00782-6</a>; PMID: <a
 href="https://www.ncbi.nlm.nih.gov/pubmed/33603233" target="_blank">33603233</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8812996/" target="_blank">PMC8812996</a>
 </p>
 
 
 
 <p>
 Kwak H, Fuda NJ, Core LJ, Lis JT.
 <a href="https://www.ncbi.nlm.nih.gov/pubmed/23430654" target="_blank">