40264d0668926b51da94d2ca038dc7349dfc3405 markd Sat Sep 26 07:59:34 2026 -0700 Shorten the Methods sections on the two TSS data pages. refs #35528 Both read like a paper's methods section rather than a track description: 43 lines over four subsections on proCapNet, 35 on encode4ProCap. Cut each to three paragraphs, what the model or assay is, how the files we serve were produced, and where the build is recorded. The subsection headings go with it. Kept every fact a user of the track needs: the window and stride, the MANE Select limit on the contribution scores, the unresolved-sequence handling on hg38, the replicate summing, the six ENCODE accessions and the file manifest. Dropped the detail that belongs to the papers, such as the DeepSHAP scalarization and the loss weighting. The container page keeps no Methods section, which is correct for a page that only points at the two data pages. diff --git src/hg/makeDb/trackDb/human/proCapNet.html src/hg/makeDb/trackDb/human/proCapNet.html index 09390d9a4b9..b6e37bcc9ea 100644 --- src/hg/makeDb/trackDb/human/proCapNet.html +++ src/hg/makeDb/trackDb/human/proCapNet.html @@ -70,75 +70,51 @@ on, and both predicted signal and contribution scores scale with it, so the y axis is not comparable between cell lines. Tracks are colored by the cell line the model was trained on:

Methods

-

Model

-

-ProCapNet adapts the BPNet architecture. It reads 2,114 bp of one-hot encoded -sequence and outputs both a base-resolution profile over the central 1,000 bp, -covering both strands as a single softmax so the model can learn strand -asymmetry, and a scalar giving the log total number of initiation events in that -window. Training used ENCODE PRO-cap read alignments in which the first read of -each pair was discarded and only the single 5'-most base of the second read was -kept, merged across replicates and kept separate by strand. Each model was -trained on all PRO-cap peaks in its cell line plus randomly sampled -DNase-hypersensitive sites from the same cell line at a 7:1 peak to background -ratio, with 7-fold cross-validation split by chromosome. Bases that are not -uniquely mappable by 36-mer reads were given zero loss weight during training. -Full details are in Cochran et al., 2024. +ProCapNet adapts the BPNet architecture: it reads 2,114 bp of sequence and +predicts a base-resolution initiation profile over the central 1,000 bp on both +strands, together with the total number of initiation events in that window. One +model was trained per cell line on ENCODE PRO-cap data, using all PRO-cap peaks +in that cell line plus sampled DNase-hypersensitive sites as background, with +7-fold cross-validation split by chromosome. Full details are in Cochran +et al., 2024.

-

Predicted PRO-cap

-

-Genome-wide predictions were generated by the Kundaje lab by applying the model -to every 2,114 bp window at a stride of 250 bp, so each base is the average of -four overlapping predictions, then averaging across the seven cross-validation -models and across the forward and reverse-complemented sequence. On hg38 no -prediction was made where most of a window was unresolved (N) in the reference. +The Kundaje lab generated the genome-wide predictions by applying each model in +2,114 bp windows at a stride of 250 bp, then averaging across the seven +cross-validation models and across the forward and reverse-complemented +sequence. On hg38 no prediction was made where most of a window was unresolved +(N) in the reference. Contribution scores were computed with DeepSHAP, which +contrasts the model's output on the real sequence against its output on 25 +dinucleotide shuffles of it, and were computed only at MANE Select transcription +start sites.

-

Sequence contribution scores

- -

-Scores were computed with DeepSHAP, which estimates each base's contribution by -contrasting the model's output on the real sequence against its output on a set -of reference sequences, here 25 dinucleotide shuffles of the sequence being -scored. Because DeepSHAP needs a single scalar to explain, the base-resolution -profile output was summarized by mean-normalizing the pre-softmax logits and -taking their dot product with the post-softmax profile, which weights each -base's logit by its predicted probability of being used and sums over the -1,000 bp output window and both strands. This is the profile or TSS-positioning -task; ProCapNet can also produce scores for its read-count task, which are not -included in this track. Each scored sequence was run through all seven cross-validation -models and in both orientations, and the scores averaged. -

- -

Source

-

The ProCapNet model implementation is at kundajelab/ProCapNet. The steps that turned the published files into these tracks are recorded in doc/$db/transcriptionStart.txt, the scripts they run are in makeDb/outside/proCapNet, and the track configuration is in trackDb/human/$db/transcriptionStart.ra.

Data Access