53ab3f1e293a91ddf71ec7568f4a6ff827bcd278 lrnassar Fri Oct 2 13:52:36 2026 -0700 Native GPN-Star track (Ye, Benegas et al., Nature 2026) on hg38, mm39, galGal6, dm6 and ce11, from the authors' Hugging Face hub. Each model is a multiWig sequence logo plus a four-allele -LLR composite using negateValues. The three hg38 models (V/M/P) sit in predictionScoresSuper via human/gpnStar.ra with /gbdb/$D bigDataUrls so -strict drops them on other human assemblies; the other four get a standalone gpnStar superTrack. The authors' bigWigs come from pyBigWig, whose bwAddIntervalSpanSteps writes the last section of each run 6 bases too long (pyBigWig #166) and makes bigWigAverageOverBed and bigWigCorrelate abort, so gpnStarRebuild.sh re-encodes them with wigToBigWig and gpnStarVerify.sh checks every per-base value is unchanged. Entropy scores and the GenArk-only Arabidopsis set are left out, and pennantIcon still has #TBD placeholders for the newsarch anchor and date. refs #38451 diff --git src/hg/makeDb/trackDb/chicken/galGal6/gpnStar.html src/hg/makeDb/trackDb/chicken/galGal6/gpnStar.html new file mode 100644 index 00000000000..eb9f1235ade --- /dev/null +++ src/hg/makeDb/trackDb/chicken/galGal6/gpnStar.html @@ -0,0 +1,150 @@ +
+This track shows scores from GPN-Star, a genomic language model that predicts how strongly selection constrains +each base, and how damaging each possible single-base substitution is likely to be. GPN-Star +learns from a whole-genome alignment of many species together with their evolutionary tree, so it +combines the cross-species conservation signal used by PhyloP and PhastCons with the surrounding +sequence context. The model shown here was trained on the alignment of 77 vertebrates to this +assembly. Scores are available for every possible single-base substitution, in coding and +non-coding sequence alike. +
++The authors evaluated the model by checking whether variants that the model predicts to be +damaging are rarer in natural populations, as expected for variants removed by selection. Among +the variants with the most damaging scores, rare variants were more enriched for GPN-Star than +for PhyloP or PhastCons computed on the same alignment. Like PhyloP and PhastCons, +GPN-Star measures evolutionary constraint rather than the effect of a variant on a particular +trait. +
+ ++The track has two parts: a set of -LLR graphs and a sequence logo. +
++-LLR graphs. There are three possible substitutions at each position, so the scores are +split across four graphs, one per alternate allele. The graph labeled "Mutation: A" shows the +score for changing the reference base to an A, and so on. At each position the graph matching +the reference base carries no substitution and is shown as zero. The score is the calibrated +log-likelihood ratio (LLR) of the alternate versus the reference allele, displayed with its sign +flipped (-LLR), so that higher bars mean a substitution the model considers more damaging. The +default view is scaled from -3 to 10; values outside that range are clipped at the edge of the +graph, and the range can be changed on the track configuration page. In dense mode the graphs are shown in +grayscale. When expanded, they use these colors: +
+| + | Positive -LLR: the alternate allele is less likely than the reference, suggesting a + constrained position or a damaging change |
|---|---|
| + | Negative -LLR: the alternate allele is more likely than the reference |
+The scores are calibrated so that a substitution behaving like those at neutral sites scores +about 0. +
++Sequence logo. At each position the logo stacks the four bases, with each letter drawn in +proportion to how strongly the model favors it. The total height of the stack shows how +constrained the position is: a tall stack, up to a maximum of 2, means the model is confident +about which base belongs there, while a flat or empty column means the position looks neutral. +The letters are always stacked in the same order, so the base the model favors most is the +tallest letter, not necessarily the top one. In tall, constrained columns this is almost always +the reference base; in short columns, where no base is strongly preferred, it often is not. +The logo is meant for viewing: it shows which positions are constrained, +but it levels off at a height of 2, so a substitution scored at -5 and one scored at -10 look +almost the same. Use the -LLR graphs to compare how damaging specific substitutions are. The letters use the standard nucleotide colors: +
+| A | |
| C | |
| G | |
| T |
+Individual scores are only visible when zoomed in to a few hundred bases. At wider zoom levels +each pixel summarizes many bases, and the graphs show the average with whiskers for the range. +
++Scores are only available on the main chromosomes. The mitochondrial genome and unplaced, +unlocalized or alternate sequences have no scores. +
+ ++GPN-Star is a transformer model trained to predict masked bases in a multiple-species +whole-genome alignment, using the species tree to weight related species by their evolutionary +distance. The LLR of a substitution is the log ratio of the predicted probabilities of the +alternate and reference alleles. To remove the effect of mutation rate, the authors subtracted +the average LLR at neutral sites, defined from PhyloP and PhastCons, with the same five-base +context. The model shown here has 85 million parameters and was trained on the UCSC +multiz77way alignment of 77 vertebrates. See Ye, Benegas et al. (2026) for details. +
++The bigWig files were downloaded from the +GPN-Star scores dataset on Hugging Face, revision +47e7f051113abab49f04f43f9107cae2cbbfd34d. We used the four signed LLR files (one per +alternate allele, with the reference allele set to zero) and the four sequence logo files. The +LLR values are the authors' mutation-rate calibrated scores, rounded to three decimal places. The +logo heights were computed by the authors from the same calibrated scores: the four base +probabilities are a softmax of the LLRs, with the reference base at zero, and each letter's +height is its probability times (2 - H), where H is the entropy in bits. The values were not +modified at UCSC; the browser flips the sign of the LLR at display time. We checked a random +sample of positions against the authors' Parquet tables to confirm that the scores line up with +the right base. The download and checking scripts are in +makeDb/scripts/gpnStar, the steps are documented in +doc/$db/gpnStar.txt, and the track settings are in +trackDb/chicken/galGal6/gpnStar.ra. +
+ ++The data can be explored interactively with the +Table Browser or the +Data Integrator. From scripts, the data can be accessed +through our API, for example +track=gpnStarLlrA for the scores for substitutions to A. +
+
+For automated download and analysis, the bigWig files can be downloaded from
+our download
+server. The files llr_A.bw to llr_T.bw hold the LLR scores and
+A.bw to T.bw the logo heights. Scores for a region can be extracted with our
+tool bigWigToBedGraph, which can be compiled from the source code or downloaded as a
+precompiled binary for your system. Instructions for downloading source code and binaries can be
+found here. For example:
+bigWigToBedGraph -chrom=chr1 -start=1000000 -end=1000100
+https://hgdownload.soe.ucsc.edu/gbdb/$db/gpnStar/llr_A.bw stdout
+
+The full-precision scores, including a separate calibrated entropy score per position that is +not shown here, are available as Parquet tables from the +GPN-Star +scores dataset. Note that those tables use one-based positions, while the bigWig files and +the Genome Browser use zero-based start coordinates. The trained models and the code are +available from the +GPN-Star collection on Hugging Face and from +GitHub. +
+ ++Thanks to Yun Song, Gonzalo Benegas, Chengzhong Ye and colleagues at the University of +California, Berkeley for making these scores available and for providing a track hub. +
+ ++Ye C, Benegas G, Albors C, Li JC, Prillo S, Fields PD, Clarke B, Song YS. + +Predicting genome-wide functional constraints with GPN-Star. +Nature. 2026 Sep 9;. +DOI: 10.1038/s41586-026-11005-5; +PMID: 42717086 +