53ab3f1e293a91ddf71ec7568f4a6ff827bcd278 lrnassar Fri Oct 2 13:52:36 2026 -0700 Native GPN-Star track (Ye, Benegas et al., Nature 2026) on hg38, mm39, galGal6, dm6 and ce11, from the authors' Hugging Face hub. Each model is a multiWig sequence logo plus a four-allele -LLR composite using negateValues. The three hg38 models (V/M/P) sit in predictionScoresSuper via human/gpnStar.ra with /gbdb/$D bigDataUrls so -strict drops them on other human assemblies; the other four get a standalone gpnStar superTrack. The authors' bigWigs come from pyBigWig, whose bwAddIntervalSpanSteps writes the last section of each run 6 bases too long (pyBigWig #166) and makes bigWigAverageOverBed and bigWigCorrelate abort, so gpnStarRebuild.sh re-encodes them with wigToBigWig and gpnStarVerify.sh checks every per-base value is unchanged. Entropy scores and the GenArk-only Arabidopsis set are left out, and pennantIcon still has #TBD placeholders for the newsarch anchor and date. refs #38451 diff --git src/hg/makeDb/trackDb/human/gpnStar.html src/hg/makeDb/trackDb/human/gpnStar.html new file mode 100644 index 00000000000..05164a6c94e --- /dev/null +++ src/hg/makeDb/trackDb/human/gpnStar.html @@ -0,0 +1,163 @@ +

Description

+

+These tracks are part of the +Deleteriousness Predictions +collection. They show scores from GPN-Star, a genomic language model that predicts how strongly +natural selection constrains each base of the genome, and how damaging each possible +single-base substitution is likely to be. GPN-Star learns from whole-genome alignments of many +species together with their evolutionary tree, so it combines the cross-species conservation +signal used by PhyloP and PhastCons with the surrounding sequence context. Scores are available +for every possible single-base substitution, in coding and non-coding sequence alike. +

+

+Three GPN-Star models were trained for the human genome, each on an alignment that covers a +different span of evolution: +

+ +

+The authors found that the best model depends on the question. In their benchmarks the +vertebrate model performed best on pathogenic missense variants (ClinVar) and on somatic cancer +variants (COSMIC). The mammal model performed best on pathogenic non-coding variants (OMIM and +HGMD) and on fine-mapped GWAS variants. The primate model captured the most complex trait +heritability. Like PhyloP and PhastCons, GPN-Star measures evolutionary constraint over a given +span of time rather than pathogenicity as such, so a highly constrained score does not by +itself mean a variant causes disease. +

+ +

Display Conventions and Configuration

+

+Each model has two tracks: a set of -LLR graphs and a sequence logo. +

+

+-LLR graphs. There are three possible substitutions at each position, so the scores are +split across four graphs, one per alternate allele. The graph labeled "Mutation: A" shows the +score for changing the reference base to an A, and so on. At each position the graph matching +the reference base carries no substitution and is shown as zero. The score is the calibrated +log-likelihood ratio (LLR) of the alternate versus the reference allele, displayed with its sign +flipped (-LLR), so that higher bars mean a substitution the model considers more damaging. The +default view is scaled from -3 to 10; values outside that range are clipped at the edge of the +graph, and the range can be changed on the track configuration page. In dense mode the graphs are shown in +grayscale. When expanded, they use these colors: +

+ + + + + +
 Positive -LLR: the alternate allele is less likely than the reference, suggesting a + constrained position or a damaging change
 Negative -LLR: the alternate allele is more likely than the reference
+

+The scores are calibrated so that a substitution behaving like those at neutral sites scores +about 0. +

+

+Sequence logo. At each position the logo stacks the four bases, with each letter drawn in +proportion to how strongly the model favors it. The total height of the stack shows how +constrained the position is: a tall stack, up to a maximum of 2, means the model is confident +about which base belongs there, while a flat or empty column means the position looks neutral. +The letters are always stacked in the same order, so the base the model favors most is the +tallest letter, not necessarily the top one. In tall, constrained columns this is almost always +the reference base; in short columns, where no base is strongly preferred, it often is not. +The logo is meant for viewing: it shows which positions are constrained, +but it levels off at a height of 2, so a substitution scored at -5 and one scored at -10 look +almost the same. Use the -LLR graphs to compare how damaging specific substitutions are. The letters use the standard nucleotide colors: +

+ + + + + +
 A
 C
 G
 T
+

+Individual scores are only visible when zoomed in to a few hundred bases. At wider zoom levels +each pixel summarizes many bases, and the graphs show the average with whiskers for the range. +

+

+Scores are only available on the main chromosomes. The mitochondrial genome and unplaced, +unlocalized or alternate sequences have no scores. +

+ +

Methods

+

+GPN-Star is a transformer model trained to predict masked bases in a multiple-species +whole-genome alignment, using the species tree to weight related species by their evolutionary +distance. The LLR of a substitution is the log ratio of the predicted probabilities of the +alternate and reference alleles. To remove the effect of mutation rate, the authors subtracted +the average LLR at neutral sites, defined from PhyloP and PhastCons, with the same five-base +context. The three human models have 200 million parameters each and were trained on the +UCSC multiz100way alignment (V), on the Zoonomia cactus447way alignment of 447 mammals (M), +and on the 243 primates in that alignment (P). See Ye, Benegas et al. (2026) for details. +

+

+The bigWig files were downloaded from the +GPN-Star scores dataset on Hugging Face, revision +47e7f051113abab49f04f43f9107cae2cbbfd34d. For each model we used the four signed LLR +files (one per alternate allele, with the reference allele set to zero) and the four sequence +logo files. The LLR values are the authors' mutation-rate calibrated scores, rounded to three +decimal places. The logo heights were computed by the authors from the same calibrated scores: +the four base probabilities are a softmax of the LLRs, with the reference base at zero, and each +letter's height is its probability times (2 - H), where H is the entropy in bits. The values were +not modified at UCSC; the browser flips the sign of the LLR at display time. We checked a random +sample of positions against the authors' Parquet tables to confirm that the scores line up with +the right base. The download and checking scripts are in +makeDb/scripts/gpnStar, the steps are documented in +doc/hg38/gpnStar.txt, and the track settings are in +trackDb/human/gpnStar.ra. +

+ +

Data Access

+

+The data can be explored interactively with the +Table Browser or the +Data Integrator. From scripts, the data can be accessed +through our API, for example +track=gpnStarLlrMA for the mammal model's scores for substitutions to A. +

+

+For automated download and analysis, the bigWig files can be downloaded from +our download +server, in one subdirectory per model (v100, m447, p243). The +files llr_A.bw to llr_T.bw hold the LLR scores and A.bw to +T.bw the logo heights. Scores for a region can be extracted with our tool +bigWigToBedGraph, which can be compiled from the source code or downloaded as a +precompiled binary for your system. Instructions for downloading source code and binaries can be +found here. For example:
+bigWigToBedGraph -chrom=chr17 -start=43044294 -end=43044394 +https://hgdownload.soe.ucsc.edu/gbdb/$db/gpnStar/m447/llr_A.bw stdout +

+

+The full-precision scores, including a separate calibrated entropy score per position that is +not shown here, are available as Parquet tables from the +GPN-Star +scores dataset. Note that those tables use one-based positions, while the bigWig files and +the Genome Browser use zero-based start coordinates. The trained models and the code are +available from the +GPN-Star collection on Hugging Face and from +GitHub. +

+ +

Credits

+

+Thanks to Yun Song, Gonzalo Benegas, Chengzhong Ye and colleagues at the University of +California, Berkeley for making these scores available and for providing a track hub. +

+ +

References

+

+Ye C, Benegas G, Albors C, Li JC, Prillo S, Fields PD, Clarke B, Song YS. + +Predicting genome-wide functional constraints with GPN-Star. +Nature. 2026 Sep 9;. +DOI: 10.1038/s41586-026-11005-5; +PMID: 42717086 +