fe4c75c272ca27de55c0b42396c47f2b00637a4a jnavarr5 Mon Sep 28 16:36:20 2026 -0700 Linking the encode4ProCap trackDb.ra, saying why PRO-cap and contribution scores are hg38 only on the container page, and limiting the ProCapNet prediction claim to the primary chromosomes, refs #35528 diff --git src/hg/makeDb/trackDb/human/encode4ProCap.html src/hg/makeDb/trackDb/human/encode4ProCap.html index 1a5f1b8a93c..e6831ba0819 100644 --- src/hg/makeDb/trackDb/human/encode4ProCap.html +++ src/hg/makeDb/trackDb/human/encode4ProCap.html @@ -1,169 +1,171 @@

Description

Transcription begins when RNA polymerase II starts making an RNA at a transcription start site. A promoter or an enhancer does not use a single start site: polymerase initiates across a spread of nearby bases, on both strands, and which bases are used is set largely by the local DNA sequence. PRO-cap is a nascent RNA run-on assay that captures the 5' end of each new transcript, so each read reports the exact base and strand of one initiation event. Unlike RNA-seq or CAGE, PRO-cap sees unstable RNAs as well as stable ones, which makes initiation at enhancers visible.

This track shows PRO-cap signal from six ENCODE 4 experiments, one per cell line, on GRCh38/hg38 only. The predictions a deep learning model makes from sequence alone, trained on this same data, are in the ProCapNet track, available on GRCh38/hg38 and T2T-CHM13/hs1.

Display Conventions and Configuration

Cell lines are listed in the table on this page, one row each. Use the Sample class facet to narrow the list, and the Experiment column to open the experiment on the ENCODE portal.

Each track is an overlay of the two strands: plus strand reads are drawn upward and minus strand reads downward. The y axis is the number of initiation events measured at that base, summed over the experiment's replicates, so tracks with deeper sequencing reach higher values. Tracks are colored by cell line:

Methods

PRO-cap, in the CoPRO form of Tome et al., 2018, runs a nuclear run-on reaction with biotinylated NTPs, captures the biotinylated nascent RNAs, selects for a 5' cap so that only transcripts carrying their original start are kept, and sequences them paired-end. The 5' end of the second read marks the base at which transcription started and its strand gives the strand of initiation. The six experiments were produced by the Yu and Lis labs at Cornell University as part of the ENCODE 4 nascent transcriptome survey described in Shah et al., 2026, and processed through the ENCODE PRO-cap pipeline.

The signal files were downloaded from the ENCODE portal, taking only the plus and minus strand signal of unique reads belonging to each experiment's default analysis. ENCODE publishes no pooled file, so the per-replicate files were summed at UCSC into one track per cell line and strand, the same merge the ProCapNet models were trained on. Total signal is conserved exactly by the summing, and minus strand signal is negative as released. The experiments are ENCSR046BCI, ENCSR100LIJ, ENCSR935RNW, ENCSR098LLB, ENCSR261KBX and ENCSR799DGV for A673, Caco-2, Calu3, HUVEC, K562 and MCF10A respectively.

The steps are recorded in doc/hg38/transcriptionStart.txt and the scripts are in +target="_blank">doc/hg38/transcriptionStart.txt, the scripts are in makeDb/outside/proCapNet, including proCapNetEncodeFiles.tsv, which records exactly which ENCODE file -accessions went into each track. +accessions went into each track, and the track configuration is in +trackDb/human/hg38/transcriptionStart.ra.

Data Access

The bigWig files are on our download server, named for the cell line, the ENCODE experiment accession and the strand, for example K562.ENCSR261KBX.pos.bw and K562.ENCSR261KBX.neg.bw.

The data can be explored interactively in table format with the Table Browser or the Data Integrator and exported from there to spreadsheet or tab-sep tables. From scripts, the data can be accessed through our API. The API returns one bigWig at a time, so name a single strand of one cell line rather than the container, for example track=encode4ProCap_K562_procap_pos.

Individual regions or the whole genome annotation can be obtained using our tool bigWigToBedGraph, which can be compiled from the source code or downloaded as a precompiled binary for your system. Instructions for downloading source code and binaries are on the utilities download page. The tool can also be used to obtain features within a given range, e.g. bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4ProCap/K562.ENCSR261KBX.pos.bw -chrom=chr21 -start=0 -end=100000000 stdout

The unmerged per-replicate files can be downloaded from the ENCODE portal under the experiment accessions listed above.

Credits

PRO-cap data was generated as part of ENCODE by Sagar Shah and the Yu and Lis labs at Cornell University. Thanks to the ENCODE Consortium and the ENCODE production laboratories for making the data available.

References

Shah SR, Chen Y, Leung AK, Navarro PVC, Paramo MI, Gupta J, Gurumurthy A, Fite RF, Weimer AK, Cochran K et al. The nascent transcriptome delineates the regulatory landscape in human health and disease. bioRxiv. 2026 Jun 9;. DOI: 10.1101/2025.09.24.676871; PMID: 41040139; PMC: PMC12485982

Tome JM, Tippens ND, Lis JT. Single-molecule nascent RNA sequencing identifies regulatory domain architecture at promoters and enhancers. Nat Genet. 2018 Nov;50(11):1533-1541. DOI: 10.1038/s41588-018-0234-5; PMID: 30349116; PMC: PMC6422046

Kwak H, Fuda NJ, Core LJ, Lis JT. Precise maps of RNA polymerase reveal how promoters direct initiation and pausing. Science. 2013 Feb 22;339(6122):950-3. DOI: 10.1126/science.1229386; PMID: 23430654; PMC: PMC3974810

Luo Y, Hitz BC, Gabdank I, Hilton JA, Kagda MS, Lam B, Myers Z, Sud P, Jou J, Lin K et al. New developments on the Encyclopedia of DNA Elements (ENCODE) data portal. Nucleic Acids Res. 2020 Jan 8;48(D1):D882-D889. DOI: 10.1093/nar/gkz1062; PMID: 31713622; PMC: PMC7061942