7e87cadb469b4e0eb4fb7f973154357cfe7fc345
lrnassar
  Mon Sep 21 15:56:18 2026 -0700
QA fixes for the mei (Mobile Insertions) track collection. refs #37524

Fix two data bugs found during QA and rebuild the affected bigBeds.
meiEul1dbToBed.py looked up samples and individuals by name, but euL1db
joins on 1-based row numbers, so neither join ever matched and the
individual count, tissues, clinical conditions and populations were empty
on all 8,991 insertions while the contributing-samples table printed row
numbers. Both loaders now key on the row number, the table prints the
sample name, and the adjacent population filter is case-insensitive so it
actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every
record, which dropped the first base of the element on the 96 GRCh38 and
111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no
anchor base; it now prefers INFO SEQ, which always matches SVLEN.

Correct seven statements on the description pages against their sources:
the HGSVC3 single-caller split was attributed to PALMER rather than
L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db
was credited with aligning the L1HS consensus when the paper says it was
processed from our RepeatMasker track, DeepMEI's network was described as
a classifier rather than a genotyper and given the wrong training set,
euL1db listed two detection methods absent from the data, and HMEID
contradicted itself on the MELT ASSESS cutoff.

Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted
callset is kept off the download server; the container page no longer
claims the whole collection is long-read, lists the two euL1db subtracks,
scopes its display conventions to the subtracks they describe, and cites
all six papers; dead and wrong track links are repointed and pinned to a
db; $db replaces hardcoded hg38 in paths on pages that serve three
assemblies; the euL1db labels no longer carry hg38 counts and a lift note
that made no sense on hg19; all six subtracks gain a dataVersion; the
euL1db filter ranges match the data; and five autoSql field descriptions
match what the files contain.

Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct
the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19
carries the two euL1db subtracks.

diff --git src/hg/makeDb/doc/hg38/mei.txt src/hg/makeDb/doc/hg38/mei.txt
index 14a96541bec..51ad3ef85d1 100644
--- src/hg/makeDb/doc/hg38/mei.txt
+++ src/hg/makeDb/doc/hg38/mei.txt
@@ -120,31 +120,31 @@
 wget http://bigdata.ibp.ac.cn/HMEID/static/download/sample_info.HMEIDv1.1.txt.gz
 
 # Convert site-level VCF (36699 MEIs) to bed9+27.
 # Source: ~/kent/src/hg/makeDb/scripts/mei/meiHmeidVcfToBed.py
 python3 ~/kent/src/hg/makeDb/scripts/mei/meiHmeidVcfToBed.py \
     MEI.GRCh38.HMEIDv1.1.vcf.gz \
     /hive/data/genomes/hg38/chrom.sizes \
     meiHmeid.bed
 # -> Read 36699 records, wrote 36699, skipped 0 + 0 + 0
 # Class distribution: Alu 26553, HERVK 126, L1 7353, SVA 2667.
 
 sort -k1,1 -k2,2n meiHmeid.bed > meiHmeid.sorted.bed
 
 bedToBigBed -tab \
     -as=$HOME/kent/src/hg/makeDb/scripts/mei/meiHmeid.as \
-    -type=bed9+27 \
+    -type=bed9+28 \
     meiHmeid.sorted.bed \
     /hive/data/genomes/hg38/chrom.sizes \
     meiHmeid.bb
 
 ############################################################
 # 2026-05-13 Claude (max) - SweGen MELT MEI callset (hg38, lifted from hg19)
 # Source: Ameur et al. 2017, Eur J Hum Genet, PMID 28832569 (SweGen cohort)
 #         Gardner et al. 2017, Genome Res, PMID 28855259 (MELT tool)
 # https://swefreq.nbis.se/dataset/SweGen
 
 # Site-level MELT v2.0.2 MEI callset on 1,000 Swedish WGS samples
 # (SweGen, Illumina HiSeq X, 150 bp PE, BWA-MEM v0.7.12 to GRCh37).
 # The VCF has no per-sample columns; INFO carries MELT_AN (allele
 # count, despite the name) and MELT_AF (allele frequency) plus the
 # usual MELT fields SVTYPE/SVLEN/TSD/ASSESS/MEIINFO/INTERNAL.
@@ -260,15 +260,74 @@
 bedToBigBed -tab \
     -as=$HOME/kent/src/hg/makeDb/scripts/mei/meiEul1db.as \
     -type=bed9+19 \
     /hive/data/genomes/hg38/bed/mei/eul1db/eul1db.hg38.sorted.bed \
     /hive/data/genomes/hg38/chrom.sizes \
     /hive/data/genomes/hg38/bed/mei/eul1db/eul1db.hg38.bb
 
 sort -k1,1 -k2,2n eul1dbRef.hg38.bed \
     > /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.sorted.bed
 bedToBigBed -tab \
     -as=$HOME/kent/src/hg/makeDb/scripts/mei/meiEul1dbRef.as \
     -type=bed9+6 \
     /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.sorted.bed \
     /hive/data/genomes/hg38/chrom.sizes \
     /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.bb
+
+############################################################
+# /gbdb symlinks for the whole collection
+# The trackDb stanza lives in trackDb/human/mei.ra and uses $D, so each
+# assembly needs its own symlink directory. The SweGen file is
+# underscore-prefixed because the callset cannot be redistributed (see
+# below), which keeps it off the public download server; the stanza also
+# carries "tableBrowser off", which blocks both hgTables and the REST API.
+
+mkdir -p /gbdb/hg38/mei /gbdb/hs1/mei /gbdb/hg19/mei
+
+ln -s /hive/data/genomes/hg38/bed/mei/meiHgsvc3.bb            /gbdb/hg38/mei/hgsvc3.bb
+ln -s /hive/data/genomes/hg38/bed/mei/deepmei/deepmei1kg.bb   /gbdb/hg38/mei/deepmei1kg.bb
+ln -s /hive/data/genomes/hg38/bed/mei/hmei/meiHmeid.bb        /gbdb/hg38/mei/hmeid.bb
+ln -s /hive/data/genomes/hg38/bed/mei/swegen/meiSwegen.bb     /gbdb/hg38/mei/_swegen.bb
+ln -s /hive/data/genomes/hg38/bed/mei/eul1db/eul1db.hg38.bb   /gbdb/hg38/mei/eul1db.bb
+ln -s /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.bb /gbdb/hg38/mei/eul1dbRef.bb
+ln -s /hive/data/genomes/hs1/bed/mei/meiHgsvc3.bb             /gbdb/hs1/mei/hgsvc3.bb
+ln -s /hive/data/genomes/hg19/bed/mei/eul1db/eul1db.hg19.bb   /gbdb/hg19/mei/eul1db.bb
+ln -s /hive/data/genomes/hg19/bed/mei/eul1db/eul1dbRef.hg19.bb /gbdb/hg19/mei/eul1dbRef.bb
+
+############################################################
+# 2026-09-21 Lou - QA fixes (refs #37524)
+
+# 1) meiHgsvc3: inserted sequence on the PALMER-only records.
+# The source CSV merges three different INFO schemas. Most records follow
+# the VCF convention (ALT[0] == REF[0], len(ALT)-1 == SVLEN), but the 96
+# GRCh38 / 111 T2T-CHM13 records whose INFO carries SEQ/CALLERS do not:
+# there ALT is the element itself (len(ALT) == SVLEN, and on 70 of the 96
+# ALT[0] differs from REF[0]), and 13 of them carry a truncated ALT.
+# meiHgsvc3CsvToBed.py took alt[1:] unconditionally, which dropped the
+# first base of the element on all 96 and left insertSeq disagreeing with
+# svLen. The script now prefers INFO SEQ when present, which always
+# matches SVLEN. Rebuilt both assemblies with the commands above;
+# item counts and every other column are unchanged.
+
+# 2) meiEul1db: sample and individual joins.
+# euL1db's foreign keys are 1-based row numbers, not names: SRIP.txt's
+# sampleID column indexes Samples.txt (943 rows) and Samples.txt's
+# Individual_id column indexes Individuals.txt (741 rows). Both totals
+# match Table 1 of Mir et al. 2015. meiEul1dbToBed.py keyed both lookup
+# tables by name, so neither join ever matched and individualCount,
+# tissues, diseases and populations were empty on all 8,991 MRIPs, while
+# the contributing-samples table printed row numbers instead of sample
+# names. Both loaders now key on the row number and the sample table
+# prints the name. The adjacent population filter compared against
+# "Unknown" while the data says "unknown", so it never fired either; it is
+# now case-insensitive. After the fix: individualCount set on 8,976
+# records, tissues 8,976, diseases 6,781, populations 7,589 (hg19; 7,586
+# after the lift to hg38, which loses 3 MRIPs). Rebuilt the
+# hg19 bigBed and re-lifted to hg38 with the commands above; item counts
+# (8,991 / 8,988) and all other columns are unchanged.
+
+# 3) meiSwegen: /gbdb/hg38/mei/swegen.bb renamed to _swegen.bb. The
+# callset cannot be redistributed, and the underscore prefix is what keeps
+# a file off the public download server. "tableBrowser off" already
+# blocked hgTables and the REST API (verified: the API returns HTTP 403
+# "protected data"), but without the prefix the bigBed itself would have
+# been mirrored to hgdownload. mei.ra was updated to match.