53ab3f1e293a91ddf71ec7568f4a6ff827bcd278 lrnassar Fri Oct 2 13:52:36 2026 -0700 Native GPN-Star track (Ye, Benegas et al., Nature 2026) on hg38, mm39, galGal6, dm6 and ce11, from the authors' Hugging Face hub. Each model is a multiWig sequence logo plus a four-allele -LLR composite using negateValues. The three hg38 models (V/M/P) sit in predictionScoresSuper via human/gpnStar.ra with /gbdb/$D bigDataUrls so -strict drops them on other human assemblies; the other four get a standalone gpnStar superTrack. The authors' bigWigs come from pyBigWig, whose bwAddIntervalSpanSteps writes the last section of each run 6 bases too long (pyBigWig #166) and makes bigWigAverageOverBed and bigWigCorrelate abort, so gpnStarRebuild.sh re-encodes them with wigToBigWig and gpnStarVerify.sh checks every per-base value is unchanged. Entropy scores and the GenArk-only Arabidopsis set are left out, and pennantIcon still has #TBD placeholders for the newsarch anchor and date. refs #38451 diff --git src/hg/makeDb/trackDb/human/gpnStar.html src/hg/makeDb/trackDb/human/gpnStar.html new file mode 100644 index 00000000000..05164a6c94e --- /dev/null +++ src/hg/makeDb/trackDb/human/gpnStar.html @@ -0,0 +1,163 @@ +<h2>Description</h2> +<p> +These tracks are part of the +<a href="hgTrackUi?g=predictionScoresSuper">Deleteriousness Predictions</a> +collection. They show scores from GPN-Star, a genomic language model that predicts how strongly +natural selection constrains each base of the genome, and how damaging each possible +single-base substitution is likely to be. GPN-Star learns from whole-genome alignments of many +species together with their evolutionary tree, so it combines the cross-species conservation +signal used by PhyloP and PhastCons with the surrounding sequence context. Scores are available +for every possible single-base substitution, in coding and non-coding sequence alike. +</p> +<p> +Three GPN-Star models were trained for the human genome, each on an alignment that covers a +different span of evolution: +</p> +<ul> +<li><b>GPN-Star (V)</b>: 100 vertebrates (about 600 million years)</li> +<li><b>GPN-Star (M)</b>: 447 mammals (about 100 million years)</li> +<li><b>GPN-Star (P)</b>: 243 primates (about 65 million years)</li> +</ul> +<p> +The authors found that the best model depends on the question. In their benchmarks the +vertebrate model performed best on pathogenic missense variants (ClinVar) and on somatic cancer +variants (COSMIC). The mammal model performed best on pathogenic non-coding variants (OMIM and +HGMD) and on fine-mapped GWAS variants. The primate model captured the most complex trait +heritability. Like PhyloP and PhastCons, GPN-Star measures evolutionary constraint over a given +span of time rather than pathogenicity as such, so a highly constrained score does not by +itself mean a variant causes disease. +</p> + +<h2>Display Conventions and Configuration</h2> +<p> +Each model has two tracks: a set of -LLR graphs and a sequence logo. +</p> +<p> +<b>-LLR graphs.</b> There are three possible substitutions at each position, so the scores are +split across four graphs, one per alternate allele. The graph labeled "Mutation: A" shows the +score for changing the reference base to an A, and so on. At each position the graph matching +the reference base carries no substitution and is shown as zero. The score is the calibrated +log-likelihood ratio (LLR) of the alternate versus the reference allele, displayed with its sign +flipped (-LLR), so that higher bars mean a substitution the model considers more damaging. The +default view is scaled from -3 to 10; values outside that range are clipped at the edge of the +graph, and the range can be changed on the track configuration page. In dense mode the graphs are shown in +grayscale. When expanded, they use these colors: +</p> +<table class="stdTbl"> + <tr><th style="background-color:#8C3C3C;width:2em"> </th> + <td>Positive -LLR: the alternate allele is less likely than the reference, suggesting a + constrained position or a damaging change</td></tr> + <tr><th style="background-color:#3C3C8C;width:2em"> </th> + <td>Negative -LLR: the alternate allele is more likely than the reference</td></tr> +</table> +<p> +The scores are calibrated so that a substitution behaving like those at neutral sites scores +about 0. +</p> +<p> +<b>Sequence logo.</b> At each position the logo stacks the four bases, with each letter drawn in +proportion to how strongly the model favors it. The total height of the stack shows how +constrained the position is: a tall stack, up to a maximum of 2, means the model is confident +about which base belongs there, while a flat or empty column means the position looks neutral. +The letters are always stacked in the same order, so the base the model favors most is the +tallest letter, not necessarily the top one. In tall, constrained columns this is almost always +the reference base; in short columns, where no base is strongly preferred, it often is not. +The logo is meant for viewing: it shows which positions are constrained, +but it levels off at a height of 2, so a substitution scored at -5 and one scored at -10 look +almost the same. Use the -LLR graphs to compare how damaging specific substitutions are. The letters use the standard nucleotide colors: +</p> +<table class="stdTbl"> + <tr><th style="background-color:#008000;width:2em"> </th><td>A</td></tr> + <tr><th style="background-color:#0000FF;width:2em"> </th><td>C</td></tr> + <tr><th style="background-color:#FFA600;width:2em"> </th><td>G</td></tr> + <tr><th style="background-color:#FF0000;width:2em"> </th><td>T</td></tr> +</table> +<p> +Individual scores are only visible when zoomed in to a few hundred bases. At wider zoom levels +each pixel summarizes many bases, and the graphs show the average with whiskers for the range. +</p> +<p> +Scores are only available on the main chromosomes. The mitochondrial genome and unplaced, +unlocalized or alternate sequences have no scores. +</p> + +<h2>Methods</h2> +<p> +GPN-Star is a transformer model trained to predict masked bases in a multiple-species +whole-genome alignment, using the species tree to weight related species by their evolutionary +distance. The LLR of a substitution is the log ratio of the predicted probabilities of the +alternate and reference alleles. To remove the effect of mutation rate, the authors subtracted +the average LLR at neutral sites, defined from PhyloP and PhastCons, with the same five-base +context. The three human models have 200 million parameters each and were trained on the +UCSC multiz100way alignment (V), on the Zoonomia cactus447way alignment of 447 mammals (M), +and on the 243 primates in that alignment (P). See Ye, Benegas <em>et al.</em> (2026) for details. +</p> +<p> +The bigWig files were downloaded from the +<a href="https://huggingface.co/datasets/songlab/gpn-star-scores/tree/main/bigwig" +target="_blank">GPN-Star scores dataset on Hugging Face</a>, revision +<tt>47e7f051113abab49f04f43f9107cae2cbbfd34d</tt>. For each model we used the four signed LLR +files (one per alternate allele, with the reference allele set to zero) and the four sequence +logo files. The LLR values are the authors' mutation-rate calibrated scores, rounded to three +decimal places. The logo heights were computed by the authors from the same calibrated scores: +the four base probabilities are a softmax of the LLRs, with the reference base at zero, and each +letter's height is its probability times (2 - H), where H is the entropy in bits. The values were +not modified at UCSC; the browser flips the sign of the LLR at display time. We checked a random +sample of positions against the authors' Parquet tables to confirm that the scores line up with +the right base. The download and checking scripts are in +<a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/gpnStar" +target="_blank">makeDb/scripts/gpnStar</a>, the steps are documented in +<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/gpnStar.txt" +target="_blank">doc/hg38/gpnStar.txt</a>, and the track settings are in +<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/gpnStar.ra" +target="_blank">trackDb/human/gpnStar.ra</a>. +</p> + +<h2>Data Access</h2> +<p> +The data can be explored interactively with the +<a href="../cgi-bin/hgTables">Table Browser</a> or the +<a href="../cgi-bin/hgIntegrator">Data Integrator</a>. From scripts, the data can be accessed +through our <a href="https://api.genome.ucsc.edu" target="_blank">API</a>, for example +track=<i>gpnStarLlrMA</i> for the mammal model's scores for substitutions to A. +</p> +<p> +For automated download and analysis, the bigWig files can be downloaded from +<a href="https://hgdownload.soe.ucsc.edu/gbdb/$db/gpnStar/" target="_blank">our download +server</a>, in one subdirectory per model (<tt>v100</tt>, <tt>m447</tt>, <tt>p243</tt>). The +files <tt>llr_A.bw</tt> to <tt>llr_T.bw</tt> hold the LLR scores and <tt>A.bw</tt> to +<tt>T.bw</tt> the logo heights. Scores for a region can be extracted with our tool +<tt>bigWigToBedGraph</tt>, which can be compiled from the source code or downloaded as a +precompiled binary for your system. Instructions for downloading source code and binaries can be +found <a href="https://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" +target="_blank">here</a>. For example:<br> +<tt>bigWigToBedGraph -chrom=chr17 -start=43044294 -end=43044394 +https://hgdownload.soe.ucsc.edu/gbdb/$db/gpnStar/m447/llr_A.bw stdout</tt> +</p> +<p> +The full-precision scores, including a separate calibrated entropy score per position that is +not shown here, are available as Parquet tables from the +<a href="https://huggingface.co/datasets/songlab/gpn-star-scores" target="_blank">GPN-Star +scores dataset</a>. Note that those tables use one-based positions, while the bigWig files and +the Genome Browser use zero-based start coordinates. The trained models and the code are +available from the +<a href="https://huggingface.co/collections/songlab/gpn-star-68c0c055acc2ee51d5c4f129" +target="_blank">GPN-Star collection on Hugging Face</a> and from +<a href="https://github.com/songlab-cal/gpn" target="_blank">GitHub</a>. +</p> + +<h2>Credits</h2> +<p> +Thanks to Yun Song, Gonzalo Benegas, Chengzhong Ye and colleagues at the University of +California, Berkeley for making these scores available and for providing a track hub. +</p> + +<h2>References</h2> +<p> +Ye C, Benegas G, Albors C, Li JC, Prillo S, Fields PD, Clarke B, Song YS. +<a href="https://www.ncbi.nlm.nih.gov/pubmed/42717086" target="_blank"> +Predicting genome-wide functional constraints with GPN-Star</a>. +<em>Nature</em>. 2026 Sep 9;. +DOI: <a href="https://doi.org/10.1038/s41586-026-11005-5" target="_blank">10.1038/s41586-026-11005-5</a>; +PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/42717086" target="_blank">42717086</a> +</p>