7485753cb34c1c3fd4a3f7b8894d993110db50ff
lrnassar
  Thu Jul 23 14:03:39 2026 -0700
Rebuild popEVE heatmap with the full per-amino-acid matrix. refs #37791

The track was built from the genomic single-nucleotide VCF, so each column showed only the
~6 of 19 substitutions reachable by a single nucleotide change. The popEVE authors confirmed
the per-transcript CSV download carries the complete matrix, so the converter now takes the
all-19 scores from those CSVs while reusing the genomic codon coordinates, strand, and blocks
derived from the VCF. Proteins absent from the (older) CSV release fall back to the sparse VCF
scores. Color anchors recomputed over the full-matrix distribution, and the mouseover's raw
EVE field is relabeled "EVE index" to distinguish it from the EVE track's 0-1 score. Updates
the description page and makedoc to match.

diff --git src/hg/makeDb/trackDb/human/popEve.html src/hg/makeDb/trackDb/human/popEve.html
index 473d820640a..9bd05a3f109 100644
--- src/hg/makeDb/trackDb/human/popEve.html
+++ src/hg/makeDb/trackDb/human/popEve.html
@@ -1,182 +1,184 @@
 <h2>Description</h2>
 <p>
 Missense mutations change a single amino acid in a protein and make up a large fraction
 of the variants observed in human populations, yet most have no established clinical
 significance. <a href="https://pop.evemodel.org/" target="_blank">popEVE</a> places missense
 variants on a single, proteome-wide spectrum of
 deleteriousness, so that variants in different genes can be compared directly. It is a deep
 generative model that combines cross-species evolutionary scores with human population
 variation: scores from EVE (an evolutionary variational autoencoder) and the ESM-1v protein
 language model are calibrated against allele observations in the UK Biobank using a Gaussian
 process, yielding a continuous, human-specific measure of variant impact. This track shows
-popEVE scores for single-nucleotide missense substitutions across roughly 18,000 human proteins.
+popEVE scores for all missense substitutions across roughly 18,000 human proteins.
 </p>
 
 <h2>Display Conventions</h2>
 <p>
 Each entry spans one protein at its genomic locus. The heatmap columns correspond to
 individual amino acid positions in the protein, placed at the codon's genomic coordinate.
 The rows correspond to the 20 standard amino acids, ordered by amino acid class to match
 the MaveDB track. Each cell
 shows the popEVE score for substituting the wildtype amino acid at that position with the
-row amino acid. Because popEVE is distributed as a list of genomic single-nucleotide
-variants, only amino acid substitutions reachable by a single-nucleotide change are scored;
-cells for substitutions requiring more than one nucleotide change, for the wildtype amino
-acid, or for positions without a score (for example start codons) are left empty.
+row amino acid. All 19 possible substitutions are scored at each position; only the wildtype
+cell is left empty. A small number of proteins, taken from an earlier data release that
+provided only genomic single-nucleotide variants, show just the substitutions reachable by a
+single-nucleotide change (roughly 6 of 19 per position) and therefore appear sparser.
 </p>
 
 <p>
 Unlike per-gene scores, popEVE is calibrated across the whole proteome, so cells are colored
 on a single global gradient keyed to the raw popEVE score (lower, more negative scores are
 more deleterious). The color is interpolated between the five anchors below: the published
 severe and moderate thresholds are fixed anchors, and the extremes saturate at the 0.5th and
 99.5th percentiles of the proteome-wide score distribution.
 </p>
 <table style="border-collapse: collapse; border: 1px solid #ccc;">
 <tr><th style="border:1px solid #ccc;padding:6px 12px;background:#f2f2f2;">Color</th>
     <th style="border:1px solid #ccc;padding:6px 12px;background:#f2f2f2;">popEVE score</th>
     <th style="border:1px solid #ccc;padding:6px 12px;background:#f2f2f2;">Interpretation</th></tr>
 <tr><td style="border:1px solid #ccc;background:#b2182b;width:50px;">&nbsp;</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">&le; &minus;5.74</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">Most deleterious (color saturates here)</td></tr>
 <tr><td style="border:1px solid #ccc;background:#d6604d;width:50px;">&nbsp;</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">&asymp; &minus;5.056</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">Severe threshold: high-confidence deleterious (99.99% likelihood of falling in the more deleterious distribution)</td></tr>
 <tr><td style="border:1px solid #ccc;background:#f4a582;width:50px;">&nbsp;</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">&asymp; &minus;4.617</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">Moderate threshold</td></tr>
 <tr><td style="border:1px solid #ccc;background:#f7f7f7;width:50px;">&nbsp;</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">&asymp; &minus;3.5</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">Near the proteome-wide median</td></tr>
 <tr><td style="border:1px solid #ccc;background:#2166ac;width:50px;">&nbsp;</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">&ge; &minus;2.29</td>
     <td style="border:1px solid #ccc;padding:4px 10px;">Most tolerated (color saturates here)</td></tr>
 </table>
 <p>
 <b>Note:</b> popEVE ranks deleteriousness to organismal fitness, weighted toward severe,
 often early-onset phenotypes, rather than classic clinical pathogenicity. Some well-known
 disease genes whose variants act mainly through loss of function or cause adult-onset
 conditions (for example BRCA1) may therefore show few or no cells in the severe range.
 </p>
 
 <p>
 Hovering over a cell shows a summary of that substitution and the scores behind it, for
 example:
 </p>
 <p style="border:1px solid #ccc; padding:6px 12px; display:inline-block;">
 G1042&rarr;A<br>
 <b>popEVE:</b> -3.652 (tolerated)<br>
-<b>EVE:</b> 6.75<br>
+<b>EVE index:</b> 6.75<br>
 <b>ESM1v:</b> -8.779<br>
 <b>popAdj:</b> EVE -3.537 / ESM1v -3.767<br>
 <b>gap:</b> 0.277
 </p>
 <table style="border-collapse: collapse; border: 1px solid #ccc;">
 <tr><th style="border:1px solid #ccc;padding:6px 12px;background:#f2f2f2;">Field</th>
     <th style="border:1px solid #ccc;padding:6px 12px;background:#f2f2f2;">Meaning</th></tr>
 <tr><td style="border:1px solid #ccc;padding:4px 10px;"><tt>G1042&rarr;A</tt></td>
     <td style="border:1px solid #ccc;padding:4px 10px;">The substitution: the wildtype amino acid (G) at the protein position (1042) changed to the variant amino acid (A) represented by this row.</td></tr>
 <tr><td style="border:1px solid #ccc;padding:4px 10px;"><tt>popEVE</tt></td>
     <td style="border:1px solid #ccc;padding:4px 10px;">The popEVE score and its severity class (severe / moderate / tolerated). This is the calibrated, proteome-wide score &mdash; comparable across genes, with lower (more negative) values more deleterious &mdash; and it determines the cell color.</td></tr>
-<tr><td style="border:1px solid #ccc;padding:4px 10px;"><tt>EVE</tt></td>
+<tr><td style="border:1px solid #ccc;padding:4px 10px;"><tt>EVE index</tt></td>
     <td style="border:1px solid #ccc;padding:4px 10px;">The raw EVE evolutionary index used as an input to popEVE (a log-likelihood-ratio scale on which higher values are more deleterious). This is not the same as the 0&ndash;1 score shown in the separate EVE track.</td></tr>
 <tr><td style="border:1px solid #ccc;padding:4px 10px;"><tt>ESM1v</tt></td>
     <td style="border:1px solid #ccc;padding:4px 10px;">The raw ESM-1v protein language model score used as the other input to popEVE, on which more negative values are more deleterious.</td></tr>
 <tr><td style="border:1px solid #ccc;padding:4px 10px;"><tt>popAdj</tt></td>
     <td style="border:1px solid #ccc;padding:4px 10px;">The EVE and ESM-1v scores after popEVE's population-based recalibration onto the proteome-wide scale; the popEVE score is the combination of these two.</td></tr>
 <tr><td style="border:1px solid #ccc;padding:4px 10px;"><tt>gap</tt></td>
     <td style="border:1px solid #ccc;padding:4px 10px;">The fraction of sequences with a gap at this position in the multiple sequence alignment (a per-position value; higher means sparser alignment coverage and a weaker evolutionary signal).</td></tr>
 </table>
 
 <p>
 For reverse-strand genes, protein positions are displayed left to right in genomic
 order (C-terminus to N-terminus on the screen), consistent with the standard genome
 browser orientation.
 </p>
 
 <p>
 popEVE scores are computed per protein isoform, so a gene with more than one RefSeq isoform
 appears as several overlapping entries (one per protein accession) that may stack in the
 display. The track item name is the gene symbol; the specific RefSeq protein accession is
 available as a field in the Table Browser and details page.
 </p>
 
 <h2>Methods</h2>
 <p>
 popEVE is built on top of two unsupervised variant effect models: EVE, a Bayesian
 variational autoencoder trained per protein on multiple sequence alignments of evolutionary
 variation, and ESM-1v, a protein language model trained on the UniRef database. For each
 protein, the evolutionary scores from these models are treated as a prior, and a Gaussian
 process is used to model how the probability of observing a variant in the human population
 (the UK Biobank, ~500,000 exomes) depends on its evolutionary score. Population data are
 used only as a coarse "seen / not seen" signal rather than by allele frequency, which lets
 the resulting score rank the deleteriousness of variants across different proteins while
 leaving the within-gene ranking essentially unchanged. The inferred function rescales the
 evolutionary score per protein to account for the degree of constraint acting on that
 protein in the population, producing a calibrated, proteome-wide score. Low (more negative)
 scores indicate greater predicted deleteriousness. Following the authors, a high-confidence
 severe threshold is set at &minus;5.056 and a moderate threshold at &minus;4.617, each
 derived from a two-component Gaussian mixture fit. See Orenbuch et al. 2025 for full details.
 </p>
 
 <p>
-The data were downloaded as a complete GRCh38 VCF (UKBB-trained popEVE, version 1.1) from
-<a href="https://pop.evemodel.org" target="_blank">https://pop.evemodel.org</a>.
-The VCF lists, for every possible missense substitution, its genomic coordinates, the RefSeq
-protein accession and gene symbol, the popEVE score and component scores. Multiple codon
-changes encoding the same amino acid substitution carry identical popEVE scores and were
-deduplicated to one record per substitution. Records were converted to heatmap bigBed format
+The full per-amino-acid popEVE score matrices (all substitutions at every position, with the
+popEVE score and component scores) were downloaded as per-transcript files from
+<a href="https://pop.evemodel.org" target="_blank">https://pop.evemodel.org</a>
+(UKBB-trained popEVE, version 1.1). Because those files are keyed by RefSeq protein position
+rather than by genome coordinate, the genomic codon coordinates for each residue were taken
+from the companion complete GRCh38 popEVE VCF. A small number of proteins present in the VCF
+but not in the per-transcript release retain only their single-nucleotide-reachable
+substitutions. Records were converted to heatmap bigBed format
 with a custom Python script; full processing instructions are in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/popEve.txt"
 target="_blank">makedoc file</a>, and the conversion scripts are available in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/popEve"
 target="_blank">our GitHub repository</a>. The strand of each protein was taken from NCBI
 RefSeq. The global color gradient saturates at the 0.5th and 99.5th percentiles of the
 proteome-wide score distribution.
 </p>
 
 <h2>Data Access</h2>
 <p>The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to
 spreadsheet or tab-separated tables. From scripts, the data can be accessed through our
 <a href="https://api.genome.ucsc.edu" target="_blank">API</a>, track=<i>popEve</i>.</p>
 <p>For automated download and analysis, the genome annotation is stored in a bigBed file
 that can be downloaded from
 <a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/popEve" target="_blank">our download
 server</a>. The file for this track is called <tt>popEve.bb</tt>. Individual regions or the
 whole genome annotation can be obtained using our tool <tt>bigBedToBed</tt>, which can be
 compiled from the source code or downloaded as a precompiled binary for your system.
 Instructions for downloading source code and binaries can be found
 <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" target="_blank">here</a>.
 The tool can also be used to obtain features within a given range, e.g.
 <tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/$db/popEve/popEve.bb -chrom=chr17
 -start=43000000 -end=43200000 stdout</tt></p>
 <p>The original annotation source data can be downloaded from
 <a href="https://pop.evemodel.org" target="_blank">https://pop.evemodel.org</a>.</p>
 
 <h2>Credits</h2>
 <p>
 Thanks to Rose Orenbuch, Mafalda Dias, Jonathan Frazer, Debora S. Marks, and colleagues at
 Harvard Medical School, the Centre for Genomic Regulation, and collaborating institutions for
 developing popEVE and making the scores publicly available at
 <a href="https://pop.evemodel.org" target="_blank">pop.evemodel.org</a>.
 </p>
 
 <h2>References</h2>
 <p>
 Orenbuch R, Shearer CA, Kollasch AW, Spinner AD, Hopf TA, van Niekerk L, Franceschi D,
 Dias M, Frazer J, Marks DS.
 <a href="https://doi.org/10.1038/s41588-025-02400-1" target="_blank">
 Proteome-wide model for human disease genetics</a>.
 <em>Nat Genet</em>. 2025 Dec;57(12):3165-3174.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/41286104" target="_blank">41286104</a>
 </p>
 
 <p>
 Frazer J, Notin P, Dias M, Gomez A, Min JK, Brock K, Gal Y, Marks DS.
 <a href="https://doi.org/10.1038/s41586-021-04043-8" target="_blank">
 Disease variant prediction with deep generative models of evolutionary data</a>.
 <em>Nature</em>. 2021 Nov;599(7883):91-95.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/34707284" target="_blank">34707284</a>
 </p>