7e87cadb469b4e0eb4fb7f973154357cfe7fc345
lrnassar
  Mon Sep 21 15:56:18 2026 -0700
QA fixes for the mei (Mobile Insertions) track collection. refs #37524

Fix two data bugs found during QA and rebuild the affected bigBeds.
meiEul1dbToBed.py looked up samples and individuals by name, but euL1db
joins on 1-based row numbers, so neither join ever matched and the
individual count, tissues, clinical conditions and populations were empty
on all 8,991 insertions while the contributing-samples table printed row
numbers. Both loaders now key on the row number, the table prints the
sample name, and the adjacent population filter is case-insensitive so it
actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every
record, which dropped the first base of the element on the 96 GRCh38 and
111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no
anchor base; it now prefers INFO SEQ, which always matches SVLEN.

Correct seven statements on the description pages against their sources:
the HGSVC3 single-caller split was attributed to PALMER rather than
L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db
was credited with aligning the L1HS consensus when the paper says it was
processed from our RepeatMasker track, DeepMEI's network was described as
a classifier rather than a genotyper and given the wrong training set,
euL1db listed two detection methods absent from the data, and HMEID
contradicted itself on the MELT ASSESS cutoff.

Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted
callset is kept off the download server; the container page no longer
claims the whole collection is long-read, lists the two euL1db subtracks,
scopes its display conventions to the subtracks they describe, and cites
all six papers; dead and wrong track links are repointed and pinned to a
db; $db replaces hardcoded hg38 in paths on pages that serve three
assemblies; the euL1db labels no longer carry hg38 counts and a lift note
that made no sense on hg19; all six subtracks gain a dataVersion; the
euL1db filter ranges match the data; and five autoSql field descriptions
match what the files contain.

Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct
the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19
carries the two euL1db subtracks.

diff --git src/hg/makeDb/scripts/mei/meiHmeid.as src/hg/makeDb/scripts/mei/meiHmeid.as
index 7e145bf5d0c..4ec59104b46 100644
--- src/hg/makeDb/scripts/mei/meiHmeid.as
+++ src/hg/makeDb/scripts/mei/meiHmeid.as
@@ -1,22 +1,22 @@
 table meiHmeid
 "Mobile Element Insertions in 5,675 genomes (HMEID v1.1 / NyuWa + 1KGP, GRCh38)"
 (
 string  chrom;             "Reference chromosome or scaffold"
 uint    chromStart;        "0-based start position (anchor base)"
 uint    chromEnd;          "Half-open end position (anchor base + 1)"
-string  name;              "Item label (INS, MEI class, carrier allele count)"
+string  name;              "Item label (element class, carrier allele count)"
 uint    score;             "Score (alt-allele frequency * 1000)"
 char[1] strand;            "Strand (always .)"
 uint    thickStart;        "Start of thick drawing region"
 uint    thickEnd;          "End of thick drawing region"
 uint    itemRgb;           "RGB color, by mobile-element class"
 string  teClass;           "TE class|Family of mobile element (Alu, L1, SVA, HERVK)"
 int     svLen;             "Insertion length (bp), -1 if unknown"
 string  tsd;               "Target site duplication|TSD sequence reported by MELT, '.' if unknown"
 int     assess;            "MELT ASSESS score|0=no overlapping reads, 1=imprecise, 2=discordant pairs only, 3=left-side TSD only, 4=right-side TSD only, 5=TSD decided with split reads (highest quality)"
 int     altAlleleCount;    "Carrier haplotypes (all)|Haplotypes carrying the insertion across all 5,675 samples"
 int     alleleNumber;      "Genotyped haplotypes (all)|Total haplotypes successfully genotyped"
 float   altAlleleFreq;     "Allele frequency (all)|Alt-allele frequency across all 5,675 samples"
 int     nyuwaAC;           "NyuWa AC|Alt-allele count in NyuWa (Chinese) cohort"
 int     nyuwaAN;           "NyuWa AN|Genotyped haplotypes in NyuWa cohort"
 float   nyuwaAF;           "NyuWa AF|Alt-allele frequency in NyuWa (Chinese) cohort"