fe4c75c272ca27de55c0b42396c47f2b00637a4a jnavarr5 Mon Sep 28 16:36:20 2026 -0700 Linking the encode4ProCap trackDb.ra, saying why PRO-cap and contribution scores are hg38 only on the container page, and limiting the ProCapNet prediction claim to the primary chromosomes, refs #35528 diff --git src/hg/makeDb/trackDb/human/proCapNet.html src/hg/makeDb/trackDb/human/proCapNet.html index 4dc96db59b2..af2aac4b48f 100644 --- src/hg/makeDb/trackDb/human/proCapNet.html +++ src/hg/makeDb/trackDb/human/proCapNet.html @@ -1,202 +1,202 @@

Description

ProCapNet is a neural network trained to predict PRO-cap signal from DNA sequence alone. PRO-cap is a run-on assay that captures the 5' end of each nascent RNA, so it reports the exact base and strand at which RNA polymerase II started transcribing, including at enhancers and at unstable transcripts that RNA-seq and CAGE miss. Six separate models were trained, one on each of six cell lines with ENCODE PRO-cap data.

This track holds two kinds of output from those models:

The predictions are available on GRCh38/hg38 and T2T-CHM13/hs1. The contribution scores are available on GRCh38/hg38 only, because they are computed at MANE Select transcription start sites and MANE is not defined for T2T-CHM13.

Predictions are not measurements: they say what the sequence looks capable of, not what a given cell is doing. The matching experimental data is in the PRO-cap track, available on GRCh38/hg38.

Display Conventions and Configuration

Cell lines are listed in the table on this page, one row each, with a checkbox per data type. Use the Sample class facet to narrow the list, and the Group tracks by buttons to order the browser by sample or by data type.

Each predicted PRO-cap track is an overlay of the two strands: plus strand predictions are drawn upward and minus strand predictions downward. The y axis is the predicted number of PRO-cap reads at that base.

Contribution scores are drawn as a sequence logo when zoomed in far enough to show individual bases: the letter of the reference base is scaled by its score, so a run of tall letters is a motif the model relied on. At lower zoom the same values are drawn as a wiggle. Scores can be negative, meaning the base argued against initiation being placed where it was. Scores exist only in the roughly 2 kb window around each MANE Select transcription start site, about 38.7 Mb of the genome. Everywhere else the track is empty, which is not the same as a score of zero.

Read depth differs between the six PRO-cap experiments the models were trained on, and both predicted signal and contribution scores scale with it, so the y axis is not comparable between cell lines. Tracks are colored by the cell line the model was trained on:

Methods

ProCapNet adapts the BPNet architecture: it reads 2,114 bp of sequence and predicts a base-resolution initiation profile over the central 1,000 bp on both strands, together with the total number of initiation events in that window. One model was trained per cell line on ENCODE PRO-cap data, using all PRO-cap peaks in that cell line plus sampled DNase-hypersensitive sites as background, with 7-fold cross-validation split by chromosome. Full details are in Cochran et al., 2024.

The Kundaje lab generated the genome-wide predictions by applying each model in 2,114 bp windows at a stride of 250 bp, then averaging across the seven cross-validation models and across the forward and reverse-complemented sequence. On hg38 no prediction was made where most of a window was unresolved (N) in the reference. Contribution scores were computed with DeepSHAP, which contrasts the model's output on the real sequence against its output on 25 dinucleotide shuffles of it, and were computed only at MANE Select transcription start sites.

The ProCapNet model implementation is at kundajelab/ProCapNet. The steps that turned the published files into these tracks are recorded in doc/$db/transcriptionStart.txt, the scripts they run are in makeDb/outside/proCapNet, and the track configuration is in trackDb/human/$db/transcriptionStart.ra.

Data Access

The bigWig files are on our download server. Predictions are under pred/ and are named for the cell line, the model and the strand, for example K562.proCapNet.pos.bw and K562.proCapNet.neg.bw. Contribution scores, which exist for GRCh38 only, are under contrib/, for example K562.proCapNet-contrib.bw.

The data can be explored interactively in table format with the Table Browser or the Data Integrator and exported from there to spreadsheet or tab-sep tables. From scripts, the data can be accessed through our API. The API returns one bigWig at a time, so name a single strand of one cell line rather than the container, for example track=proCapNet_K562_pred_pos.

Individual regions or the whole genome annotation can be obtained using our tool bigWigToBedGraph, which can be compiled from the source code or downloaded as a precompiled binary for your system. Instructions for downloading source code and binaries are on the utilities download page. The tool can also be used to obtain features within a given range, e.g. bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/$db/proCapNet/pred/K562.proCapNet.pos.bw -chrom=chr21 -start=0 -end=100000000 stdout

The ProCapNet models are on the ENCODE portal as BPNet-model annotations, one per cell line, linked from the Experiment column of the table on this page. Each annotation also holds the trained model, sequence contribution scores and predicted signal over a selected set of regions. The genome-wide predictions shown here are not part of that ENCODE release.

Credits

ProCapNet was developed by Kelly Cochran in the Kundaje lab at Stanford University. The genome-wide predictions and contribution scores were generated by Kelly Cochran in collaboration with the GENCODE consortium. Thanks to Kelly Cochran and Anshul Kundaje for making the data available.

References

Cochran K, Yin M, Mantripragada A, Schreiber J, Marinov GK, Shah SR, Yu H, Lis JT, Kundaje A. Dissecting the cis-regulatory syntax of transcription initiation with deep learning. bioRxiv. 2024 Nov 21;. DOI: 10.1101/2024.05.28.596138; PMID: 38853896; PMC: PMC11160661

Avsec Ž, Weilert M, Shrikumar A, Krueger S, Alexandari A, Dalal K, Fropf R, McAnany C, Gagneur J, Kundaje A et al. Base-resolution models of transcription-factor binding reveal soft motif syntax. Nat Genet. 2021 Mar;53(3):354-366. DOI: 10.1038/s41588-021-00782-6; PMID: 33603233; PMC: PMC8812996

Kwak H, Fuda NJ, Core LJ, Lis JT. Precise maps of RNA polymerase reveal how promoters direct initiation and pausing. Science. 2013 Feb 22;339(6122):950-3. DOI: 10.1126/science.1229386; PMID: 23430654; PMC: PMC3974810