7e87cadb469b4e0eb4fb7f973154357cfe7fc345
lrnassar
  Mon Sep 21 15:56:18 2026 -0700
QA fixes for the mei (Mobile Insertions) track collection. refs #37524

Fix two data bugs found during QA and rebuild the affected bigBeds.
meiEul1dbToBed.py looked up samples and individuals by name, but euL1db
joins on 1-based row numbers, so neither join ever matched and the
individual count, tissues, clinical conditions and populations were empty
on all 8,991 insertions while the contributing-samples table printed row
numbers. Both loaders now key on the row number, the table prints the
sample name, and the adjacent population filter is case-insensitive so it
actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every
record, which dropped the first base of the element on the 96 GRCh38 and
111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no
anchor base; it now prefers INFO SEQ, which always matches SVLEN.

Correct seven statements on the description pages against their sources:
the HGSVC3 single-caller split was attributed to PALMER rather than
L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db
was credited with aligning the L1HS consensus when the paper says it was
processed from our RepeatMasker track, DeepMEI's network was described as
a classifier rather than a genotyper and given the wrong training set,
euL1db listed two detection methods absent from the data, and HMEID
contradicted itself on the MELT ASSESS cutoff.

Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted
callset is kept off the download server; the container page no longer
claims the whole collection is long-read, lists the two euL1db subtracks,
scopes its display conventions to the subtracks they describe, and cites
all six papers; dead and wrong track links are repointed and pinned to a
db; $db replaces hardcoded hg38 in paths on pages that serve three
assemblies; the euL1db labels no longer carry hg38 counts and a lift note
that made no sense on hg19; all six subtracks gain a dataVersion; the
euL1db filter ranges match the data; and five autoSql field descriptions
match what the files contain.

Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct
the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19
carries the two euL1db subtracks.

diff --git src/hg/makeDb/trackDb/human/meiHmeid.html src/hg/makeDb/trackDb/human/meiHmeid.html
index 183d6b9cf7b..ea7e49f6d05 100644
--- src/hg/makeDb/trackDb/human/meiHmeid.html
+++ src/hg/makeDb/trackDb/human/meiHmeid.html
@@ -61,81 +61,81 @@
 from split reads).
 </p>
 
 <h2>Methods</h2>
 <p>
 HMEID was built by Niu et al. (2022) from Illumina short-read whole-genome
 sequencing of two cohorts: 2,999 individuals from the NyuWa dataset
 (diabetes and control samples collected across China, median depth
 ~26.2&times; on GRCh38) and 2,691 samples from the 1000 Genomes Project
 (~7.4&times; coverage, GRCh38-aligned CRAMs from EBI). Non-reference
 MEIs were detected with MELT v2.1.5 in SPLIT mode with default
 parameters; BAM coverage was estimated with goleft v0.1.8 covstats.
 After the MELT MakeVCF step, sites were filtered to those that (i) lie
 outside low-complexity regions, (ii) are genotyped in &gt;25% of
 individuals, (iii) have more than 2 split reads, (iv) carry a MELT
-ASSESS score &gt;3 (i.e. ASSESS &ge; 4 in the unfiltered output, but
-HMEID retains ASSESS 3 sites that otherwise pass) and (v) are marked
-PASS in the FILTER column. Alu and L1 subfamilies were assigned by
+ASSESS score of at least 3, meaning target-site duplication evidence on
+at least one side, and (v) are marked PASS in the FILTER column. Alu and L1 subfamilies were assigned by
 MELT's CALU and LINEU modules. 2,998 of 2,999 NyuWa samples and 2,677
 of 2,691 1KGP samples passed processing, yielding the final callset of
 36,699 MEIs in 5,675 genomes. Allele frequencies were computed per
 cohort and per 1KGP super-population with BCFtools v1.3.1. See
 Niu et al. 2022 for full methodological details.
 </p>
 
 <p>
 The site-level VCF was downloaded from the
 <a href="http://bigdata.ibp.ac.cn/HMEID/download/" target="_blank">HMEID
 download page</a> (file
 <tt>MEI.GRCh38.HMEIDv1.1.vcf.gz</tt>) and converted to bigBed following
 the steps in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/mei.txt"
 target="_blank">makeDoc file</a>. Conversion uses scripts in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/mei"
-target="_blank">src/hg/makeDb/scripts/mei</a>: VCF-style positions
+target="_blank">src/hg/makeDb/scripts/mei</a>,
+and the track configuration is in <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/mei.ra" target="_blank">trackDb/human/mei.ra</a>: VCF-style positions
 (1-based POS, anchor base) are converted to half-open BED coordinates
 (<tt>chromStart = POS - 1</tt>, <tt>chromEnd = chromStart + 1</tt>),
 MELT SVTYPE codes (<tt>ALU</tt>, <tt>LINE1</tt>, <tt>SVA</tt>,
 <tt>HERVK</tt>) are mapped to the element class names used here, and
 INFO fields are copied through to per-cohort and per-super-population
 allele count / number / frequency columns. All 36,699 input records
 produced one BED row each; no records were dropped.
 </p>
 
 <h2>Data Access</h2>
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from
 there to spreadsheet or tab-separated tables. From scripts, the data can
-be accessed through our <a href="https://api.genome.ucsc.edu">API</a>,
+be accessed through our <a href="https://api.genome.ucsc.edu" target="_blank">API</a>,
 track=<i>meiHmeid</i>.
 </p>
 <p>
 For automated download and analysis, the genome annotation is stored in
 a bigBed file that can be downloaded from
-<a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/" target="_blank">
+<a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/" target="_blank">
 our download server</a>. The file for this track is called
-<tt>hmeid.bb</tt> in <tt>/gbdb/hg38/mei/</tt>. Individual regions or the
+<tt>hmeid.bb</tt> in <tt>/gbdb/$db/mei/</tt>. Individual regions or the
 whole genome annotation can be obtained using our tool
 <tt>bigBedToBed</tt>, which can be compiled from the source code or
 downloaded as a precompiled binary for your system. Instructions for
 downloading source code and binaries can be found
-<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">here</a>.
+<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" target="_blank">here</a>.
 The tool can also be used to obtain features within a given range, e.g.
-<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/hmeid.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>.
+<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/hmeid.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>.
 </p>
 <p>
 The original annotation source data can be downloaded from the
 <a href="http://bigdata.ibp.ac.cn/HMEID/download/" target="_blank">HMEID
 download page</a>.
 </p>
 
 <h2>Credits</h2>
 <p>
 Thanks to Yiwei Niu, Shunmin He and colleagues at the Institute of
 Biophysics, Chinese Academy of Sciences for building HMEID and
 releasing the callset, and to the NyuWa project and the 1000 Genomes
 Project for producing the underlying whole-genome sequencing data.
 </p>