116aaf68aa340197a2a63856a90f533909e15f92
lrnassar
  Fri Oct 2 14:25:53 2026 -0700
BRCAmlaZanti.py: match ENIGMA PP4/BP5 variants by normalized genomic allele instead of HGVS name, so the 297 variants the source papers spell differently (c.4574_4575del vs Li's c.4574_4575delAA, Zanti's ins for a dup, del15) become one item with their LRs multiplied, as Finja Hennig asked. Every item is now drawn the way the ClinVar track draws it (deleted bases shifted left, 2 bp flank for ins/dup), names drop spelled-out bases, LRs are written with 5 significant figures, and the Zanti <CNV> row is dropped. Zanti rows are keyed from their own VCF columns because hgvsToVcf mis-converts intronic insertions (#38469), and the pre-Zanti input now comes from archive/v1.1 since /gbdb BRCAmfa.bb is this script's own output. The hgvsToVcf FILTER and del/dup count checks were added after a Claude review. refs #38467

diff --git src/hg/makeDb/doc/enigma.txt src/hg/makeDb/doc/enigma.txt
index 9766849413f..6cf0c1cd635 100644
--- src/hg/makeDb/doc/enigma.txt
+++ src/hg/makeDb/doc/enigma.txt
@@ -1,104 +1,149 @@
 #RM#32919
 
 mkdir /hive/data/inside/enigmaTracksData
 # excel data provided by Anna on RM and converted to txt and uploaded to directory for all tracks
 mkdir /gbdb/hg38/bbi/enigma
 mkdir /gbdb/hg19/bbi/enigma
 
 #The 5 tracks were then created by individual scripts that can all be found in the following directory:
 ~/kent/src/hg/makeDb/scripts/enigma/
 
 #Quick link for github: https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/enigma
 
 #############################################################################
 # Update to CSpec specification v1.2 (2026-08-17) RM #38130
 
 # ClinGen released v1.2 of the ENIGMA BRCA1/BRCA2 specifications (approved
 # 2025-01-09): BRCA1 GN092 (doi 10.5281/zenodo.21434315), BRCA2 GN097
 # (doi 10.5281/zenodo.21434343). Comparison against the v1.1 tables showed data
 # changes only in Table 4 (splice-site PVS1 codes) and Table 9 (PMIDs and typo
 # fixes); ST1 exon weights and the clinical domain definitions are unchanged, so
 # only BRCAsplicing and BRCAfunctionalAssays were rebuilt. BRCAmla is built from
 # publications and is independent of the specification version.
 
 mkdir /hive/data/inside/enigmaTracksData/v1.2
 # Source files downloaded from the CSpec registry "Files & Images" panel
 # (https://cspec.genome.network/cspec/ui/svi/doc/GN092):
 # Table 4:  https://cspec.genome.network/cspec/File/id/ca5cf57b-94df-4ad6-a001-c62ceccb3845/data
 # Table 9:  https://cspec.genome.network/cspec/File/id/0a35d6a8-5050-44b6-8a9d-babe8cdc06b2/data
 # SuppTbls: https://cspec.genome.network/cspec/File/id/cb4a09fe-30f4-4aa8-9d76-d7ea407c9754/data
 # Spec doc: https://cspec.genome.network/cspec/File/id/11e62fec-23b0-4a3e-b2df-751855301746/data
 # saved as CSpec_BRCA12ACMG_Rules-Specifications_V1.2_Table-4.xlsx etc.
 
 # Export the needed sheets to text. Merged cells are expanded; the Table 9
 # banner row v1.2 inserted is dropped so the layout matches the v1.1 export.
 # The new "Dace & Findlay, Interim Report" sheet in the Table 9 xlsx holds
 # interim (uncalibrated) results and is intentionally not used.
 python3 ~/kent/src/hg/makeDb/scripts/enigma/exportV12Sheets.py
 
 # The v1.2 Table 4 excel is a visual per-exon layout, unlike the flat table used
 # for v1.1, so a converter rebuilds the flat 8-column format the track script
 # consumes. Exons split by the NMD-escape boundary are encoded in v1.2 as
 # PTC<p.X / PTC>p.Y qualifiers; the converter turns those back into the c. sub-
 # ranges used in v1.1. Text is kept as UTF-8 (v1.1 text had mangled the Greek
 # delta to "?").
 python3 ~/kent/src/hg/makeDb/scripts/enigma/convertTable4toFlat.py
 
 # Rebuild the two tracks. Both scripts now write into the v1.2/ dir; the hub and
 # the /gbdb symlinks keep pointing at the fixed filenames one level up, which are
 # only overwritten at release (below). The two haplotype variants in Table 9
 # (c.[5359T>A;5363G>A] and c.[1073T>G;1078T>C;1084G>C;1086G>T]) cannot be
 # converted by hgvsToVcf and are skipped, same as in the v1.1 build.
 python3 ~/kent/src/hg/makeDb/scripts/enigma/BRCAfunctionalAssays.py
 python3 ~/kent/src/hg/makeDb/scripts/enigma/BRCAsplicing.py
 
 # Release: copy the verified .bb files onto the staging filenames the symlink
 # chain serves (do NOT touch the symlinks themselves), then copy the updated
 # hub.txt, trackDb.txt, enigma.html and the v1.2 raw files into
 # /hive/data/outside/enigma/ (= htdocs-hgdownload/hubs/enigma).
 # for db in Hg19 Hg38; do for t in BRCAsplicing BRCAfunctionalAssays; do
 #   cp /hive/data/inside/enigmaTracksData/v1.2/$t$db.bb /hive/data/inside/enigmaTracksData/$t$db.bb.tmp
 #   mv /hive/data/inside/enigmaTracksData/$t$db.bb.tmp /hive/data/inside/enigmaTracksData/$t$db.bb
 # done; done
 
 #############################################################################
 # BRCAmla: add Zanti et al. 2025 case-control LRs (2026-08-18) RM #37886
 
 # At the request of the ENIGMA collaborators, the case-control component of the
 # PP4/BP5 multifactorial likelihood track was updated from the iCOGS-derived
 # values in Parsons et al. 2019 (20 variants) to the case-control likelihood
 # ratios (ccLR) from Zanti et al. 2025 (Nat Commun, PMID 40413188,
 # doi 10.1038/s41467-025-59979-6), a case-control analysis of the BRIDGES,
 # CARRIERS and UK Biobank cohorts. The old iCOGS values were dropped rather
 # than kept alongside because iCOGS overlaps the Zanti cohorts (all 20 variants
 # recur in the Zanti data) and keeping both would count the same evidence twice.
 
 mkdir /hive/data/inside/enigmaTracksData/zantiDraft
 # Supplementary Data 4 of the paper saved there as ZantiSuppData4.xlsx
 # (also copied to /hive/data/outside/enigma/rawData/ at release).
 
 # The build script reads the current BRCAmfa bigBeds for both assemblies to
 # reuse the existing family-history, co-occurrence, segregation and pathology
 # LRs and their coordinates, drops the old case-control column, and merges in
 # the Zanti ccLR keyed on transcript:HGVSc. The new combined LR is the product
 # of the available evidence types. Variant universe is the union of the current
 # track and the Zanti variants with a computable ccLR (Zanti rows with
 # suggested code N/A or no ccLR are skipped). The new .as adds per-cohort
 # columns (BRIDGES, CARRIERS, UK Biobank) and Zanti's standalone suggested
 # code; output is bed9+17.
 python3 ~/kent/src/hg/makeDb/scripts/enigma/BRCAmlaZanti.py
 # Result: 13,481 variants per assembly (up from 4,436), written as
 # BRCAmfaZantiHg38.bb / BRCAmfaZantiHg19.bb in the zantiDraft dir. The script
 # also writes directionConflicts.tsv listing the 180 variants where the prior
 # multifactorial evidence and the ccLR point in opposite directions; these are
 # multiplied through as usual per collaborator consensus (Andreas Laner et al.,
 # see RM #37886) and a caveat was added to the hub description page.
 
 # Release, same procedure as the v1.2 update above: copy the verified .bb onto
 # the staging filenames the /gbdb symlink chain serves (symlinks untouched),
 # then the updated enigma.html and trackDb.txt (dataVersion line added, type
 # corrected from bed9+67 to bed9+17) into /hive/data/outside/enigma/.
 # for db in Hg19 Hg38; do
 #   cp /hive/data/inside/enigmaTracksData/zantiDraft/BRCAmfaZanti$db.bb /hive/data/inside/enigmaTracksData/BRCAmfa$db.bb.new
 #   mv /hive/data/inside/enigmaTracksData/BRCAmfa$db.bb.new /hive/data/inside/enigmaTracksData/BRCAmfa$db.bb
 # done
+
+##############################################################################
+# PP4/BP5 track: merge equivalent variant names, ClinVar-style positions
+# (DONE - 2026-10-02 - Lou, refs #38467)
+# The source studies spell some variants differently (Li et al. writes the
+# deleted bases out, c.4574_4575delAA; Parsons, Caputo and Zanti write
+# c.4574_4575del; Zanti writes some dups as ins), and the #37886 build matched
+# Zanti to the older track by name string, so 297 variants appeared as two or
+# three items. BRCAmlaZanti.py was revised to:
+#  - key every variant by a left-normalized genomic (chrom, pos, ref, alt):
+#    pre-Zanti items from their HGVS names via hgvsToVcf, Zanti rows from
+#    Zanti's own CHR/POS/REF/ALT columns. hgvsToVcf mis-converts insertions
+#    with an intronic offset (refs #38469), which is why Zanti rows are not
+#    keyed from their HGVSc. 7 Zanti rows carry a REF that does not match the
+#    genome (all A>G) and are keyed from their HGVSc instead.
+#  - multiply the per-type LRs (family history, co-occurrence, segregation,
+#    pathology) and the per-source LR lists of all items sharing a key. Li et
+#    al. lists c.815_824dup twice (39 and 1 probands); Li multiplies
+#    per-proband LRs, so the two rows are multiplied.
+#  - write names in current HGVS style (spelled-out bases and counts dropped;
+#    ins that duplicates adjacent sequence renamed to the dup that vcfToHgvs
+#    reports).
+#  - draw every item the way the ClinVar track does: SNV on its base, deletion
+#    on the deleted bases shifted left, delins on the replaced bases,
+#    insertion/dup on the 2 bases flanking the left-shifted insertion point.
+#    hg19 positions are computed the same way, not lifted.
+#  - drop the Zanti row with symbolic alleles (<CNV>).
+#  - write LRs with 5 significant figures instead of 5 decimals.
+# The pre-Zanti input is now read from archive/v1.1, since /gbdb BRCAmfa.bb
+# has served this script's own output since the #37886 release.
+mkdir /hive/data/inside/enigmaTracksData/rm38467
+cd /hive/data/inside/enigmaTracksData/rm38467
+python3 ~/kent/src/hg/makeDb/scripts/enigma/BRCAmlaZanti.py
+# Result: 13,182 variants per assembly (13,481 before; 595 items merged into
+# 297, one <CNV> row dropped), written as BRCAmfaZantiHg38.bb /
+# BRCAmfaZantiHg19.bb. mergedGroups.tsv lists the merged items, renamed.tsv
+# the 479 renamed singletons, directionConflicts.tsv the 210 variants whose
+# older evidence and ccLR point in opposite directions.
+
+# Release, same procedure as above:
+# for db in Hg19 Hg38; do
+#   cp /hive/data/inside/enigmaTracksData/rm38467/BRCAmfaZanti$db.bb /hive/data/inside/enigmaTracksData/BRCAmfa$db.bb.new
+#   mv /hive/data/inside/enigmaTracksData/BRCAmfa$db.bb.new /hive/data/inside/enigmaTracksData/BRCAmfa$db.bb
+# done
+# and enigma.html (Changelog and Methods text) into /hive/data/outside/enigma/.