7713a08da69691ba499d5b9d44379c43e8cdb608 markd Fri Sep 18 09:27:52 2026 -0700 Adding Transcription Start container with ENCODE 4 PRO-cap and ProCapNet tracks. refs #35528 New superTrack transcriptionStart in the rna group, holding two faceted composites: encode4ProCap with PRO-cap measurements and proCapNet with the model predictions and sequence-contribution scores. hg38 has all three data types, hs1 the predictions only. ENCODE 4 PRO-cap comes from the portal rather than the submitter's hub copy. For each of the six experiments only the plus and minus strand signal of unique reads files of that experiment's default analysis are taken, which drops files superseded by a later reprocessing. ENCODE publishes no pooled file, so the per-replicate files are summed per cell line and strand, the same merge the ProCapNet models were trained on. Total signal is conserved exactly. The published ProCapNet prediction bigWigs store one bedGraph interval per base and hold a literal NaN at every unresolved (N) base on hg38, which makes autoScale and every summary statistic NaN. They are re-encoded into fixedStep sections, a third smaller with no value changed, dropping 164,268,582 NaN bases of 3,088,269,832 on hg38 and none on hs1. Losslessness verified against the originals on random windows across five chromosomes. The composites are faceted rather than plain because a container multiWig under a plain composite is flattened away by hgTrackUi and never drawn. Each cell line is one row with a checkbox per data type, a Sample class facet, ENCODE accession links and a Files column linking each bigWig on hgdownload. Scripts and the cell line configuration are in makeDb/outside/proCapNet; the trackDb stanzas and the faceted metadata tables are generated, not hand edited. Claude-Session: https://claude.ai/code/session_01LAB6jWshLvB7eNXQKWVuW5 diff --git src/hg/makeDb/doc/hs1/transcriptionStart.txt src/hg/makeDb/doc/hs1/transcriptionStart.txt new file mode 100644 index 00000000000..4c4be06e33e --- /dev/null +++ src/hg/makeDb/doc/hs1/transcriptionStart.txt @@ -0,0 +1,42 @@ +# Transcription start sites: ProCapNet predictions, #35528, Claude +Thu Sep 17 2026 (Claude/markd) + +# hs1 has the ProCapNet predictions only. There is no PRO-cap experiment and no +# sequence-contribution score set on this assembly. See +# makeDb/doc/hg38/transcriptionStart.txt for the hg38 build, which carries all +# three. Scripts are in ~/kent/src/hg/makeDb/outside/proCapNet. + +mkdir -p /hive/data/outside/proCapNet/hs1 /hive/data/genomes/hs1/bed/proCapNet +cd /hive/data/genomes/hs1/bed/proCapNet +~/kent/src/hg/makeDb/outside/proCapNet/proCapNetDownload hs1 \ + ~/kent/src/hg/makeDb/outside/proCapNet/proCapNetExperiments.tsv \ + /hive/data/outside/proCapNet/hs1/pred +# 300 GB, about 30 minutes at six parallel streams + +# Re-encode from one bedGraph interval per base into fixedStep sections. Unlike +# hg38 these files hold no NaN, so every base is kept. +~/kent/src/hg/makeDb/outside/proCapNet/proCapNetPredBuild \ + /hive/data/genomes/hs1/chrom.sizes /hive/data/outside/proCapNet/hs1/pred pred 6 +# all 12 files report: +# 24 chroms, 3117275501 bases, 3117275501 with data, 0 dropped +# about 25 GB in, 16.7 to 17.0 GB out per file + +# the mean matches the published file exactly, so nothing was altered +bigWigInfo pred/K562.proCapNet.pos.bw | egrep 'basesCovered|mean' +# basesCovered: 3,117,275,501 +# mean: 0.019773 + +cd ~/kent/src/hg/makeDb/outside/proCapNet +./proCapNetTrackDb hs1 proCapNetExperiments.tsv \ + ~/kent/src/hg/makeDb/trackDb/human/hs1/transcriptionStart.ra /gbdb/hs1 + +# transcriptionStart.ra is generated; the include line in human/hs1/trackDb.ra +# is added by hand: +# include transcriptionStart.ra alpha + +# hs1 is a curated hub assembly, so the stanzas reach the browser through the +# hub built under /gbdb/hs1/hubs/$USER by the trackDb make, and are visible only +# when curatedHubPrefix in the sandbox hg.conf names that directory. + +# the downloads are only needed for the re-encoding +rm -rf /hive/data/outside/proCapNet/hs1/pred