29af14b47208040441ab4ff7acf51aac71cad4e2 lrnassar Thu Sep 17 21:58:18 2026 -0700 Adding MaveMD track, clinically calibrated multiplexed variant effect measurements on hg38. refs #37800 MaveMD is a curated collection inside MaveDB holding the score sets with clinical relevance, restricted to genes with a moderate or stronger gene-disease association. Data come from the MaveDB API rather than the Zenodo snapshots, at the request of the MaveDB team, who also set the request pacing the fetch script follows. Two views under a new phenDis container. mavemdVar draws one item per measured variant per score set, carrying the assay score, the functional class, the ACMG/AMP functional evidence code and strength where the score set is calibrated, the assay metadata, and matching ClinVar and gnomAD annotations. mavemdMap draws each score set as a variant effect map, one column per amino acid position and one row per substitution, reusing Jonathan's heatmap display from the MaveDB track. MaveDB resolves about a third of the collection to the genome and the rest only to a protein sequence, so variants are placed from the genomic HGVS term where one exists, otherwise by projecting the protein term onto its codon through ncbiRefSeqLink and ncbiRefSeqCurated for RefSeq accessions or the GENCODE tables for Ensembl ones, otherwise by running the submitted transcript term through hgvsToVcf. The projection is cross-checked against the variants carrying both coordinate systems and the build aborts if they disagree beyond a threshold. Codons split across an exon junction are written as BED12 with the two real blocks. Haplotypes cannot be given a single position and are excluded, with every dropped measurement counted by reason in the makeDoc. Also adds reciprocal relatedTracks entries between mavedb and mavemd. Gated alpha pending QA. diff --git src/hg/makeDb/scripts/mavemd/runBuild.sh src/hg/makeDb/scripts/mavemd/runBuild.sh new file mode 100755 index 00000000000..c5d0dda0624 --- /dev/null +++ src/hg/makeDb/scripts/mavemd/runBuild.sh @@ -0,0 +1,43 @@ +#!/bin/bash +# Build both MaveMD tracks for hg38 from a fetchMaveMd.py download. +# +# runBuild.sh +# +# Produces, in : mavemdVar.bb (one item per variant per score set) and +# mavemdMap.bb (one variant effect map per score set), plus mavemdFilters.ra, the +# generated trackDb filterValues block, and a log per stage. +set -e +set -o pipefail +export PATH=$PATH:$HOME/bin/x86_64 + +DOWNLOAD=${1:?usage: runBuild.sh } +BUILD=${2:?usage: runBuild.sh } +SCRIPTS=$(cd "$(dirname "$0")" && pwd) +CHROMSIZES=/hive/data/genomes/hg38/chrom.sizes + +mkdir -p "$BUILD" +cd "$BUILD" + +echo "[$(date +%T)] 1. per-variant bed" +"$SCRIPTS/makeMaveMdVariants.py" "$DOWNLOAD" mavemdVar.bed \ + --raFragment mavemdFilters.ra --workDir . 2> buildVar.log +tail -n 14 buildVar.log + +echo "[$(date +%T)] 2. heatmap bed" +"$SCRIPTS/makeMaveMdHeatmap.py" "$DOWNLOAD" mavemdMap.bed 2> buildMap.log +tail -n 9 buildMap.log + +echo "[$(date +%T)] 3. sort" +bedSort mavemdVar.bed mavemdVar.sorted.bed +bedSort mavemdMap.bed mavemdMap.sorted.bed + +echo "[$(date +%T)] 4. bigBed" +bedToBigBed -type=bed12+34 -tab -as="$SCRIPTS/mavemdVariants.as" \ + -extraIndex=name,clinGenId,variantUrn \ + mavemdVar.sorted.bed "$CHROMSIZES" mavemdVar.bb 2>&1 | tail -2 +bedToBigBed -type=bed12+20 -tab -as="$SCRIPTS/mavemdHeatmap.as" \ + mavemdMap.sorted.bed "$CHROMSIZES" mavemdMap.bb 2>&1 | tail -2 + +echo "[$(date +%T)] done" +bigBedInfo mavemdVar.bb | egrep 'itemCount|basesCovered|fieldCount' +bigBedInfo mavemdMap.bb | egrep 'itemCount|basesCovered|fieldCount'