7713a08da69691ba499d5b9d44379c43e8cdb608 markd Fri Sep 18 09:27:52 2026 -0700 Adding Transcription Start container with ENCODE 4 PRO-cap and ProCapNet tracks. refs #35528 New superTrack transcriptionStart in the rna group, holding two faceted composites: encode4ProCap with PRO-cap measurements and proCapNet with the model predictions and sequence-contribution scores. hg38 has all three data types, hs1 the predictions only. ENCODE 4 PRO-cap comes from the portal rather than the submitter's hub copy. For each of the six experiments only the plus and minus strand signal of unique reads files of that experiment's default analysis are taken, which drops files superseded by a later reprocessing. ENCODE publishes no pooled file, so the per-replicate files are summed per cell line and strand, the same merge the ProCapNet models were trained on. Total signal is conserved exactly. The published ProCapNet prediction bigWigs store one bedGraph interval per base and hold a literal NaN at every unresolved (N) base on hg38, which makes autoScale and every summary statistic NaN. They are re-encoded into fixedStep sections, a third smaller with no value changed, dropping 164,268,582 NaN bases of 3,088,269,832 on hg38 and none on hs1. Losslessness verified against the originals on random windows across five chromosomes. The composites are faceted rather than plain because a container multiWig under a plain composite is flattened away by hgTrackUi and never drawn. Each cell line is one row with a checkbox per data type, a Sample class facet, ENCODE accession links and a Files column linking each bigWig on hgdownload. Scripts and the cell line configuration are in makeDb/outside/proCapNet; the trackDb stanzas and the faceted metadata tables are generated, not hand edited. Claude-Session: https://claude.ai/code/session_01LAB6jWshLvB7eNXQKWVuW5 diff --git src/hg/makeDb/trackDb/human/transcriptionStart.html src/hg/makeDb/trackDb/human/transcriptionStart.html new file mode 100644 index 00000000000..5506f6625a7 --- /dev/null +++ src/hg/makeDb/trackDb/human/transcriptionStart.html @@ -0,0 +1,100 @@ +
+Transcription begins when RNA polymerase II starts making an RNA at a +transcription start site (TSS). A gene or a regulatory element does not use a +single start site: polymerase initiates across a spread of nearby bases, on both +strands, and which bases get used is set largely by the local DNA sequence and by +the chromatin state around it. Knowing where transcription starts, and how +sharply, matters for reading a promoter, for telling a promoter from an enhancer, +and for deciding which of a gene's annotated transcripts is actually made in a +given cell. +
+ ++Transcription start site is the common term and the one used here, but it names +a position where transcription initiation is the event. Initiation is the more +accurate word: polymerase does not start at one fixed point but initiates over a +window of bases, at rates that differ base by base, so what these tracks report +is the distribution of initiation across a region rather than a single site. +
+ ++No single assay settles where a TSS is. Run-on methods catch the nascent RNA at +the moment it is made, cap-based methods read the protected 5' end of a finished +transcript, long reads follow a transcript from one end to the other, and +sequence models predict initiation without an experiment at all. Each sees a +different slice of the same event, and they disagree in informative ways. This +collection gathers those lines of evidence in one place so they can be compared +at a locus. +
+ ++Tracks currently in the collection: +
+ ++Each track has its own description page covering what it measures or predicts, +how it was made, and how to download it. On T2T-CHM13 only the ProCapNet +predictions are available, since the PRO-cap experiments and the contribution +scores were produced against GRCh38. More TSS evidence tracks will be added here +over time. +
+ ++Tracks are grouped by the evidence they carry, with measurements before +predictions. Signal tracks that distinguish the two strands draw the plus strand +upward and the minus strand downward, so a divergent promoter reads as a pair of +peaks straddling the element. +
+ ++Each track listed above has its own description page with details on methods, +file locations and how to download and intersect the data. +
+ ++PRO-cap data was generated as part of ENCODE by Sagar Shah and the Yu and Lis +labs at Cornell University. ProCapNet was developed by Kelly Cochran in the +Kundaje lab at Stanford University, and the genome-wide predictions and +contribution scores were generated in collaboration with the GENCODE consortium. +
+ ++Cochran K, Yin M, Mantripragada A, Schreiber J, Marinov GK, Shah SR, Yu H, Lis JT, Kundaje A. + +Dissecting the cis-regulatory syntax of transcription initiation with deep learning. +bioRxiv. 2024 Nov 21;. +PMID: 38853896; PMC: PMC11160661 +
+ + + ++Kwak H, Fuda NJ, Core LJ, Lis JT. + +Precise maps of RNA polymerase reveal how promoters direct initiation and pausing. +Science. 2013 Feb 22;339(6122):950-3. +PMID: 23430654; PMC: PMC3974810 +
+