66fafe88184bb3a8a45a50026f4d6c7fbbc24109 markd Tue Sep 29 10:12:53 2026 -0700 Rebuild the TSS tracks as traditional composites, for wiggle control. refs #35528 A faceted composite is routed to facetedCompositeUi(), which returns before cfgByCfgType(), so it never draws the wiggle controls: no data view scaling, no viewing range, no windowing function, no track height. Signal tracks need those, so proCapNet and encode4ProCap are now traditional composites. They stay separate composites under the TSS container. The multiWig strand overlays survive the change. The old comment here claimed a multiWig under a plain composite "is flattened away and never drawn"; that is wrong for drawing, which #36320 fixed. Only the hgTrackUi subtrack list flattens, because compositeUiSubtracks() walks to leaves, so the overlays get no inline config block and are configured from their own pages. Raised as #38441. The matrix has to be declared over the leaves, since that is the level hgTrackDb checks: declaring it on the containers fails -strict with "has groups not defined in parent". So a strand is a matrix cell, and the containers carry no subGroups. Sample class survives as a filterComposite dimension, replacing the facet. configurable on gives each subtrack its own config namespace. Drop the faceted machinery: the metadata table, the color file, the constants feeding them and the gbdbDir argument, 39 lines. The metadata.tsv and colors.json already written under /gbdb are now unreferenced and can be deleted. Put the ENCODE accession in each subtrack longLabel, and add a linked table of them to both description pages. A longLabel cannot carry a link, since printSubtrackTableBody() htmlEncodes it, hence the table. The label wording is shortened to keep the longest at 74 characters. diff --git src/hg/makeDb/trackDb/human/proCapNet.html src/hg/makeDb/trackDb/human/proCapNet.html index af2aac4b48f..0ff87ecd610 100644 --- src/hg/makeDb/trackDb/human/proCapNet.html +++ src/hg/makeDb/trackDb/human/proCapNet.html @@ -31,33 +31,34 @@

The predictions are available on GRCh38/hg38 and T2T-CHM13/hs1. The contribution scores are available on GRCh38/hg38 only, because they are computed at MANE Select transcription start sites and MANE is not defined for T2T-CHM13.

Predictions are not measurements: they say what the sequence looks capable of, not what a given cell is doing. The matching experimental data is in the PRO-cap track, available on GRCh38/hg38.

Display Conventions and Configuration

-Cell lines are listed in the table on this page, one row each, with a checkbox -per data type. Use the Sample class facet to narrow the list, and the -Group tracks by buttons to order the browser by sample or by data type. +The matrix on this page has one row per cell line and one column per data type, +so a checkbox turns on one strand of one cell line's predictions, or its +contribution scores. Use the Sample class filter to restrict the matrix to +cancer or non-cancer lines.

Each predicted PRO-cap track is an overlay of the two strands: plus strand predictions are drawn upward and minus strand predictions downward. The y axis is the predicted number of PRO-cap reads at that base.

Contribution scores are drawn as a sequence logo when zoomed in far enough to show individual bases: the letter of the reference base is scaled by its score, so a run of tall letters is a motif the model relied on. At lower zoom the same values are drawn as a wiggle. Scores can be negative, meaning the base argued against initiation being placed where it was. Scores exist only in the roughly 2 kb window around each MANE Select transcription start site, about 38.7 Mb of @@ -68,30 +69,45 @@

Read depth differs between the six PRO-cap experiments the models were trained on, and both predicted signal and contribution scores scale with it, so the y axis is not comparable between cell lines. Tracks are colored by the cell line the model was trained on:

+

+One model was trained per cell line. Each links to its ENCODE annotation, +which holds the trained model itself. +

+ + + + + + + + + +
Cell lineTissueSample classENCODE model
A673MuscleCancerENCSR072YCM
Caco-2ColonCancerENCSR182QNJ
Calu3LungCancerENCSR797DEF
HUVECBlood vesselNon-cancerENCSR801ECP
K562BloodCancerENCSR740IPL
MCF10ABreastNon-cancerENCSR860TYZ
+

Methods

ProCapNet adapts the BPNet architecture: it reads 2,114 bp of sequence and predicts a base-resolution initiation profile over the central 1,000 bp on both strands, together with the total number of initiation events in that window. One model was trained per cell line on ENCODE PRO-cap data, using all PRO-cap peaks in that cell line plus sampled DNase-hypersensitive sites as background, with 7-fold cross-validation split by chromosome. Full details are in Cochran et al., 2024.

The Kundaje lab generated the genome-wide predictions by applying each model in 2,114 bp windows at a stride of 250 bp, then averaging across the seven