7e87cadb469b4e0eb4fb7f973154357cfe7fc345 lrnassar Mon Sep 21 15:56:18 2026 -0700 QA fixes for the mei (Mobile Insertions) track collection. refs #37524 Fix two data bugs found during QA and rebuild the affected bigBeds. meiEul1dbToBed.py looked up samples and individuals by name, but euL1db joins on 1-based row numbers, so neither join ever matched and the individual count, tissues, clinical conditions and populations were empty on all 8,991 insertions while the contributing-samples table printed row numbers. Both loaders now key on the row number, the table prints the sample name, and the adjacent population filter is case-insensitive so it actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every record, which dropped the first base of the element on the 96 GRCh38 and 111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no anchor base; it now prefers INFO SEQ, which always matches SVLEN. Correct seven statements on the description pages against their sources: the HGSVC3 single-caller split was attributed to PALMER rather than L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db was credited with aligning the L1HS consensus when the paper says it was processed from our RepeatMasker track, DeepMEI's network was described as a classifier rather than a genotyper and given the wrong training set, euL1db listed two detection methods absent from the data, and HMEID contradicted itself on the MELT ASSESS cutoff. Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted callset is kept off the download server; the container page no longer claims the whole collection is long-read, lists the two euL1db subtracks, scopes its display conventions to the subtracks they describe, and cites all six papers; dead and wrong track links are repointed and pinned to a db; $db replaces hardcoded hg38 in paths on pages that serve three assemblies; the euL1db labels no longer carry hg38 counts and a lift note that made no sense on hg19; all six subtracks gain a dataVersion; the euL1db filter ranges match the data; and five autoSql field descriptions match what the files contain. Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19 carries the two euL1db subtracks. diff --git src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py index ddc5f351ae6..310200c8502 100755 --- src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py +++ src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py @@ -140,31 +140,39 @@ carrierSamples.append(sampleName) altAF = (altAC / AN) if AN > 0 else 0.0 score = max(0, min(1000, int(round(altAF * 1000)))) cls = shortClass(teDesignation) color = COLOR_BY_CLASS.get(cls, COLOR_BY_CLASS["Other"]) refSd = float(info.get("REF_SD", 0)) if info.get("REF_SD", False) else 0.0 refTrf = "True" if info.get("REF_TRF", False) else "False" sourceSample = info.get("SAMPLE", "") callerCount = int(row[idx["Caller_Count"]]) if row[idx["Caller_Count"]] else 0 l1meAid = "Yes" if row[idx["L1ME-AID"]] == "1" else "No" palmer = "Yes" if row[idx["PALMER"]] == "1" else "No" - # Inserted DNA: VCF convention puts the anchor base at ALT[0] (= REF[0]). + # Inserted DNA. Most records follow the VCF convention of carrying the + # anchor base at ALT[0] (= REF[0]), so the element is ALT minus that + # base. The PALMER-only records do not: ALT there is the element + # itself (len(ALT) == SVLEN, and ALT[0] usually differs from REF[0]), + # and some of them carry a truncated ALT. Those records supply the + # full element in INFO SEQ, which always matches SVLEN, so prefer it. + if "SEQ" in info and info["SEQ"] is not True: + insertSeq = info["SEQ"] + else: insertSeq = alt[1:] if len(alt) > 1 else "" # Name format: -: (e.g. Alu-281:33). name = f"{cls}-{svLen}:{carrierCount}" out.write("\t".join([ chrom, str(chromStart), str(chromEnd), name, str(score), ".", str(chromStart), str(chromEnd), color,