29af14b47208040441ab4ff7acf51aac71cad4e2
lrnassar
  Thu Sep 17 21:58:18 2026 -0700
Adding MaveMD track, clinically calibrated multiplexed variant effect measurements on hg38. refs #37800

MaveMD is a curated collection inside MaveDB holding the score sets with clinical
relevance, restricted to genes with a moderate or stronger gene-disease association.
Data come from the MaveDB API rather than the Zenodo snapshots, at the request of the
MaveDB team, who also set the request pacing the fetch script follows.

Two views under a new phenDis container. mavemdVar draws one item per measured variant
per score set, carrying the assay score, the functional class, the ACMG/AMP functional
evidence code and strength where the score set is calibrated, the assay metadata, and
matching ClinVar and gnomAD annotations. mavemdMap draws each score set as a variant
effect map, one column per amino acid position and one row per substitution, reusing
Jonathan's heatmap display from the MaveDB track.

MaveDB resolves about a third of the collection to the genome and the rest only to a
protein sequence, so variants are placed from the genomic HGVS term where one exists,
otherwise by projecting the protein term onto its codon through ncbiRefSeqLink and
ncbiRefSeqCurated for RefSeq accessions or the GENCODE tables for Ensembl ones,
otherwise by running the submitted transcript term through hgvsToVcf. The projection is
cross-checked against the variants carrying both coordinate systems and the build aborts
if they disagree beyond a threshold. Codons split across an exon junction are written as
BED12 with the two real blocks. Haplotypes cannot be given a single position and are
excluded, with every dropped measurement counted by reason in the makeDoc.

Also adds reciprocal relatedTracks entries between mavedb and mavemd. Gated alpha
pending QA.

diff --git src/hg/makeDb/scripts/mavemd/runBuild.sh src/hg/makeDb/scripts/mavemd/runBuild.sh
new file mode 100755
index 00000000000..c5d0dda0624
--- /dev/null
+++ src/hg/makeDb/scripts/mavemd/runBuild.sh
@@ -0,0 +1,43 @@
+#!/bin/bash
+# Build both MaveMD tracks for hg38 from a fetchMaveMd.py download.
+#
+#   runBuild.sh <downloadDir> <buildDir>
+#
+# Produces, in <buildDir>: mavemdVar.bb (one item per variant per score set) and
+# mavemdMap.bb (one variant effect map per score set), plus mavemdFilters.ra, the
+# generated trackDb filterValues block, and a log per stage.
+set -e
+set -o pipefail
+export PATH=$PATH:$HOME/bin/x86_64
+
+DOWNLOAD=${1:?usage: runBuild.sh <downloadDir> <buildDir>}
+BUILD=${2:?usage: runBuild.sh <downloadDir> <buildDir>}
+SCRIPTS=$(cd "$(dirname "$0")" && pwd)
+CHROMSIZES=/hive/data/genomes/hg38/chrom.sizes
+
+mkdir -p "$BUILD"
+cd "$BUILD"
+
+echo "[$(date +%T)] 1. per-variant bed"
+"$SCRIPTS/makeMaveMdVariants.py" "$DOWNLOAD" mavemdVar.bed \
+    --raFragment mavemdFilters.ra --workDir . 2> buildVar.log
+tail -n 14 buildVar.log
+
+echo "[$(date +%T)] 2. heatmap bed"
+"$SCRIPTS/makeMaveMdHeatmap.py" "$DOWNLOAD" mavemdMap.bed 2> buildMap.log
+tail -n 9 buildMap.log
+
+echo "[$(date +%T)] 3. sort"
+bedSort mavemdVar.bed mavemdVar.sorted.bed
+bedSort mavemdMap.bed mavemdMap.sorted.bed
+
+echo "[$(date +%T)] 4. bigBed"
+bedToBigBed -type=bed12+34 -tab -as="$SCRIPTS/mavemdVariants.as" \
+    -extraIndex=name,clinGenId,variantUrn \
+    mavemdVar.sorted.bed "$CHROMSIZES" mavemdVar.bb 2>&1 | tail -2
+bedToBigBed -type=bed12+20 -tab -as="$SCRIPTS/mavemdHeatmap.as" \
+    mavemdMap.sorted.bed "$CHROMSIZES" mavemdMap.bb 2>&1 | tail -2
+
+echo "[$(date +%T)] done"
+bigBedInfo mavemdVar.bb | egrep 'itemCount|basesCovered|fieldCount'
+bigBedInfo mavemdMap.bb | egrep 'itemCount|basesCovered|fieldCount'