fcf788d7357f6f207f2bb0b61acde9dc0507d89f
lrnassar
  Mon Sep 21 04:04:10 2026 -0700
Fix two stale figures and guard the accession interpolation in the MaveMD build, per CR. refs #37800

The makeDoc script listing still said the variant converter writes bed12+31. It writes
bed12+34, as the autoSql, mavemd.ra, runBuild.sh and the makeDoc's own build-results
line all already said.

mavemdLib.py's docstring claimed the codon projection is cross-checked against "~154k
variants that carry both a genomic and a protein term". 154,895 is the number placed by
the genomic route; the cross-check set is the smaller number that also has a resolvable
protein term. The docstring now describes the set rather than quoting a figure that
drifts with every build.

Also adds checkAccession() and calls it before either query that interpolates an
accession into SQL. Nothing can currently reach those queries with a quote in it, since
the accessions come from PROTEIN_TERM whose character class excludes one, but the regex
is a hundred lines from the query and a later edit to it should not be able to open this
up silently. Output is byte-identical to the previous build.

diff --git src/hg/makeDb/doc/hg38/mavemd.txt src/hg/makeDb/doc/hg38/mavemd.txt
index 3201e1d233a..c2b3da58f48 100644
--- src/hg/makeDb/doc/hg38/mavemd.txt
+++ src/hg/makeDb/doc/hg38/mavemd.txt
@@ -40,31 +40,31 @@
 #   Collection MaveMD: 84 score sets, modified 2026-04-30
 #   Downloaded 242.1 MB in 257 seconds
 
 # Build both tracks: the per-variant bigBed and the per-score-set heatmap.
 ~/kent/src/hg/makeDb/scripts/mavemd/runBuild.sh \
     /hive/data/outside/mavemd/2026-09-17 \
     /hive/data/genomes/hg38/bed/mavemd/2026-09-17
 
 # Scripts (all in ~/kent/src/hg/makeDb/scripts/mavemd/):
 #   fetchMaveMd.py         download the collection from the API
 #
 # GENCODE VERSION: mavemdLib.py hardcodes GENCODE_ATTRS/GENCODE_GENEPRED to the V50 tables,
 # used only for the six ENSP-stated score sets. Bump both when hg38 moves to a newer GENCODE;
 # the RefSeq route (ncbiRefSeqLink/ncbiRefSeqCurated) is unversioned and needs no attention.
 #   mavemdLib.py           shared: codon projection, color tables, BED field hygiene
-#   makeMaveMdVariants.py  per-variant bed12+31 -> mavemdVar.bb
+#   makeMaveMdVariants.py  per-variant bed12+34 -> mavemdVar.bb
 #   makeMaveMdHeatmap.py   per-score-set bed12+20 -> mavemdMap.bb
 #   mavemdVariants.as, mavemdHeatmap.as
 #   runBuild.sh            driver for the two converters plus bedToBigBed
 
 # COORDINATES. MaveDB's VRS mapper resolves about a third of the collection to the genome and
 # the rest only to a protein sequence, so placement takes three routes:
 #   154,895  MaveDB's own genomic HGVS term, run through kent's hgvsToVcf
 #   296,516  projected from the protein term onto its codon
 #     9,079  the submitter's original transcript term, run through hgvsToVcf
 #   460,490  placed, of 481,717 measurements (95.6%)
 #
 # The codon projection resolves the protein accession to its transcript and walks that CDS to
 # the requested codon. MaveDB states protein terms against RefSeq for most score sets and
 # Ensembl for six, so both routes exist: NP_ through ncbiRefSeqLink + ncbiRefSeqCurated, ENSP
 # through wgEncodeGencodeAttrsV50 + wgEncodeGencodeCompV50. All 40 protein accessions resolve.