fcf788d7357f6f207f2bb0b61acde9dc0507d89f lrnassar Mon Sep 21 04:04:10 2026 -0700 Fix two stale figures and guard the accession interpolation in the MaveMD build, per CR. refs #37800 The makeDoc script listing still said the variant converter writes bed12+31. It writes bed12+34, as the autoSql, mavemd.ra, runBuild.sh and the makeDoc's own build-results line all already said. mavemdLib.py's docstring claimed the codon projection is cross-checked against "~154k variants that carry both a genomic and a protein term". 154,895 is the number placed by the genomic route; the cross-check set is the smaller number that also has a resolvable protein term. The docstring now describes the set rather than quoting a figure that drifts with every build. Also adds checkAccession() and calls it before either query that interpolates an accession into SQL. Nothing can currently reach those queries with a quote in it, since the accessions come from PROTEIN_TERM whose character class excludes one, but the regex is a hundred lines from the query and a later edit to it should not be able to open this up silently. Output is byte-identical to the previous build. diff --git src/hg/makeDb/doc/hg38/mavemd.txt src/hg/makeDb/doc/hg38/mavemd.txt index 3201e1d233a..c2b3da58f48 100644 --- src/hg/makeDb/doc/hg38/mavemd.txt +++ src/hg/makeDb/doc/hg38/mavemd.txt @@ -40,31 +40,31 @@ # Collection MaveMD: 84 score sets, modified 2026-04-30 # Downloaded 242.1 MB in 257 seconds # Build both tracks: the per-variant bigBed and the per-score-set heatmap. ~/kent/src/hg/makeDb/scripts/mavemd/runBuild.sh \ /hive/data/outside/mavemd/2026-09-17 \ /hive/data/genomes/hg38/bed/mavemd/2026-09-17 # Scripts (all in ~/kent/src/hg/makeDb/scripts/mavemd/): # fetchMaveMd.py download the collection from the API # # GENCODE VERSION: mavemdLib.py hardcodes GENCODE_ATTRS/GENCODE_GENEPRED to the V50 tables, # used only for the six ENSP-stated score sets. Bump both when hg38 moves to a newer GENCODE; # the RefSeq route (ncbiRefSeqLink/ncbiRefSeqCurated) is unversioned and needs no attention. # mavemdLib.py shared: codon projection, color tables, BED field hygiene -# makeMaveMdVariants.py per-variant bed12+31 -> mavemdVar.bb +# makeMaveMdVariants.py per-variant bed12+34 -> mavemdVar.bb # makeMaveMdHeatmap.py per-score-set bed12+20 -> mavemdMap.bb # mavemdVariants.as, mavemdHeatmap.as # runBuild.sh driver for the two converters plus bedToBigBed # COORDINATES. MaveDB's VRS mapper resolves about a third of the collection to the genome and # the rest only to a protein sequence, so placement takes three routes: # 154,895 MaveDB's own genomic HGVS term, run through kent's hgvsToVcf # 296,516 projected from the protein term onto its codon # 9,079 the submitter's original transcript term, run through hgvsToVcf # 460,490 placed, of 481,717 measurements (95.6%) # # The codon projection resolves the protein accession to its transcript and walks that CDS to # the requested codon. MaveDB states protein terms against RefSeq for most score sets and # Ensembl for six, so both routes exist: NP_ through ncbiRefSeqLink + ncbiRefSeqCurated, ENSP # through wgEncodeGencodeAttrsV50 + wgEncodeGencodeCompV50. All 40 protein accessions resolve.