7e87cadb469b4e0eb4fb7f973154357cfe7fc345
lrnassar
  Mon Sep 21 15:56:18 2026 -0700
QA fixes for the mei (Mobile Insertions) track collection. refs #37524

Fix two data bugs found during QA and rebuild the affected bigBeds.
meiEul1dbToBed.py looked up samples and individuals by name, but euL1db
joins on 1-based row numbers, so neither join ever matched and the
individual count, tissues, clinical conditions and populations were empty
on all 8,991 insertions while the contributing-samples table printed row
numbers. Both loaders now key on the row number, the table prints the
sample name, and the adjacent population filter is case-insensitive so it
actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every
record, which dropped the first base of the element on the 96 GRCh38 and
111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no
anchor base; it now prefers INFO SEQ, which always matches SVLEN.

Correct seven statements on the description pages against their sources:
the HGSVC3 single-caller split was attributed to PALMER rather than
L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db
was credited with aligning the L1HS consensus when the paper says it was
processed from our RepeatMasker track, DeepMEI's network was described as
a classifier rather than a genotyper and given the wrong training set,
euL1db listed two detection methods absent from the data, and HMEID
contradicted itself on the MELT ASSESS cutoff.

Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted
callset is kept off the download server; the container page no longer
claims the whole collection is long-read, lists the two euL1db subtracks,
scopes its display conventions to the subtracks they describe, and cites
all six papers; dead and wrong track links are repointed and pinned to a
db; $db replaces hardcoded hg38 in paths on pages that serve three
assemblies; the euL1db labels no longer carry hg38 counts and a lift note
that made no sense on hg19; all six subtracks gain a dataVersion; the
euL1db filter ranges match the data; and five autoSql field descriptions
match what the files contain.

Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct
the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19
carries the two euL1db subtracks.

diff --git src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py
index ddc5f351ae6..310200c8502 100755
--- src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py
+++ src/hg/makeDb/scripts/mei/meiHgsvc3CsvToBed.py
@@ -140,31 +140,39 @@
                     carrierSamples.append(sampleName)
 
             altAF = (altAC / AN) if AN > 0 else 0.0
             score = max(0, min(1000, int(round(altAF * 1000))))
 
             cls = shortClass(teDesignation)
             color = COLOR_BY_CLASS.get(cls, COLOR_BY_CLASS["Other"])
 
             refSd = float(info.get("REF_SD", 0)) if info.get("REF_SD", False) else 0.0
             refTrf = "True" if info.get("REF_TRF", False) else "False"
             sourceSample = info.get("SAMPLE", "")
             callerCount = int(row[idx["Caller_Count"]]) if row[idx["Caller_Count"]] else 0
             l1meAid = "Yes" if row[idx["L1ME-AID"]] == "1" else "No"
             palmer = "Yes" if row[idx["PALMER"]] == "1" else "No"
 
-            # Inserted DNA: VCF convention puts the anchor base at ALT[0] (= REF[0]).
+            # Inserted DNA. Most records follow the VCF convention of carrying the
+            # anchor base at ALT[0] (= REF[0]), so the element is ALT minus that
+            # base. The PALMER-only records do not: ALT there is the element
+            # itself (len(ALT) == SVLEN, and ALT[0] usually differs from REF[0]),
+            # and some of them carry a truncated ALT. Those records supply the
+            # full element in INFO SEQ, which always matches SVLEN, so prefer it.
+            if "SEQ" in info and info["SEQ"] is not True:
+                insertSeq = info["SEQ"]
+            else:
                 insertSeq = alt[1:] if len(alt) > 1 else ""
 
             # Name format: <class>-<svLen>:<carrierCount> (e.g. Alu-281:33).
             name = f"{cls}-{svLen}:{carrierCount}"
 
             out.write("\t".join([
                 chrom,
                 str(chromStart),
                 str(chromEnd),
                 name,
                 str(score),
                 ".",
                 str(chromStart),
                 str(chromEnd),
                 color,