7e87cadb469b4e0eb4fb7f973154357cfe7fc345
lrnassar
  Mon Sep 21 15:56:18 2026 -0700
QA fixes for the mei (Mobile Insertions) track collection. refs #37524

Fix two data bugs found during QA and rebuild the affected bigBeds.
meiEul1dbToBed.py looked up samples and individuals by name, but euL1db
joins on 1-based row numbers, so neither join ever matched and the
individual count, tissues, clinical conditions and populations were empty
on all 8,991 insertions while the contributing-samples table printed row
numbers. Both loaders now key on the row number, the table prints the
sample name, and the adjacent population filter is case-insensitive so it
actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every
record, which dropped the first base of the element on the 96 GRCh38 and
111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no
anchor base; it now prefers INFO SEQ, which always matches SVLEN.

Correct seven statements on the description pages against their sources:
the HGSVC3 single-caller split was attributed to PALMER rather than
L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db
was credited with aligning the L1HS consensus when the paper says it was
processed from our RepeatMasker track, DeepMEI's network was described as
a classifier rather than a genotyper and given the wrong training set,
euL1db listed two detection methods absent from the data, and HMEID
contradicted itself on the MELT ASSESS cutoff.

Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted
callset is kept off the download server; the container page no longer
claims the whole collection is long-read, lists the two euL1db subtracks,
scopes its display conventions to the subtracks they describe, and cites
all six papers; dead and wrong track links are repointed and pinned to a
db; $db replaces hardcoded hg38 in paths on pages that serve three
assemblies; the euL1db labels no longer carry hg38 counts and a lift note
that made no sense on hg19; all six subtracks gain a dataVersion; the
euL1db filter ranges match the data; and five autoSql field descriptions
match what the files contain.

Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct
the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19
carries the two euL1db subtracks.

diff --git src/hg/makeDb/trackDb/human/meiHgsvc3.html src/hg/makeDb/trackDb/human/meiHgsvc3.html
index e1711b8e13b..94230b9fc1d 100644
--- src/hg/makeDb/trackDb/human/meiHgsvc3.html
+++ src/hg/makeDb/trackDb/human/meiHgsvc3.html
@@ -33,31 +33,31 @@
 of the inserted mobile element (the ALT allele minus the anchor base).
 </p>
 
 <h2>Display Conventions and Configuration</h2>
 <p>
 An insertion has zero length on the reference: it attaches between two
 adjacent reference bases without replacing any of them. Following the
 VCF convention used by the underlying PAV calls and by the other
 long-read SV tracks, each MEI is drawn as a <b>1-bp block sitting on the
 anchor base</b> &mdash; the reference base immediately to the left of
 the insertion attachment point. The inserted mobile element itself is
 not present in the reference and is therefore not drawn; its length is
 reported on the detail page (svLen) and in the item label
 (<tt>class-svLen:carrierCount</tt>). bigBed does support truly zero-width
 features between nucleotides, but for consistency with the
-<a href="hgTrackUi?g=longReadVariants">long-read SV tracks</a>, this track uses the
+<a href="hgTrackUi?db=hg38&g=longReadVariants">long-read SV tracks</a>, this track uses the
 1-bp anchor representation instead.
 </p>
 <p>
 Items are colored by element class:
 </p>
 <ul>
   <li><span style="display:inline-block;background-color:#0072B2;width:18px;height:12px;vertical-align:middle;"></span> <b>Alu</b> &mdash; SINE (Short INterspersed Element)</li>
   <li><span style="display:inline-block;background-color:#D55E00;width:18px;height:12px;vertical-align:middle;"></span> <b>L1</b> &mdash; LINE-1 (Long INterspersed Element-1)</li>
   <li><span style="display:inline-block;background-color:#009E73;width:18px;height:12px;vertical-align:middle;"></span> <b>SVA</b> (SINE-VNTR-Alu) &mdash; composite retrotransposon</li>
   <li><span style="display:inline-block;background-color:#CC79A7;width:18px;height:12px;vertical-align:middle;"></span> <b>HERVK</b> (Human Endogenous Retrovirus K) &mdash; endogenous retrovirus</li>
   <li><span style="display:inline-block;background-color:#000000;width:18px;height:12px;vertical-align:middle;"></span> <b>snRNA</b> &mdash; small nuclear RNA</li>
 </ul>
 <p>
 The score column encodes the alt-allele frequency on a 0-1000 scale.
 Filters allow restricting to specific element classes, length ranges,
@@ -65,85 +65,86 @@
 and reference repeat overlap.
 </p>
 
 <h2>Methods</h2>
 <p>
 The HGSVC3 study sequenced and de novo assembled 65 individuals
 (30 males, 35 females) representing five continental groups and 28
 populations: 30 of African, 9 of Admixed American, 8 of European, 10 of
 East Asian and 8 of South Asian descent, with three parent-child trios
 included. Each sample was sequenced to ~47-fold coverage of PacBio HiFi
 and ~56-fold coverage of Oxford Nanopore long reads (~36-fold
 ultra-long), and complemented with Strand-seq, Bionano optical mapping,
 Hi-C, Iso-Seq and RNA-seq. Haplotype-resolved diploid assemblies were
 produced and structural variants called with PAV. Mobile element
 insertions were identified from the union of two independent MEI
-callsets, L1ME-AID and PALMER; all single-caller calls were manually
+callsets, L1ME-AID and PALMER2; all single-caller calls were manually
 curated. Orthogonal validation against an independent MELT-LRA callset
-showed an average concordance of 90.8% on GRCh38. Roughly 93% of MEIs
-are supported by both callers; the remaining single-caller calls split
-about 6:1 in favour of PALMER (PALMER-only ~6%, L1ME-AID-only ~1%).
+showed a concordance of 92.5% on GRCh38 and 92.1% on T2T-CHM13. Roughly 93% of MEIs
+are supported by both callers (93.6% on GRCh38, 92.4% on T2T-CHM13); the
+remaining single-caller calls split about 6:1 in favour of L1ME-AID
+(L1ME-AID-only 5.5% on GRCh38, PALMER-only 0.9%).
 The Caller Count, PALMER Validated and L1ME-AID Validated filters can
 be used to restrict the display to the high-confidence dual-validated
 subset. Calls are restricted to non-low-confidence regions
 (i.e. excluding Yq12 and centromeres).
 For each site, per-sample genotypes from all 65 assembled samples are
 summarized into an alt-allele count, allele number, allele frequency
 and a list of carrier samples. See Logsdon et al. 2025 (Nature) for
 full methodological details.
 </p>
 
 <p>
 The original CSVs were downloaded from the
 <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/HGSVC3/release/Mobile_Elements/1.0/"
 target="_blank">HGSVC3 Mobile Elements release directory</a>
 (files <tt>MEI_Callset_GRCh38.ALL.20241211.csv.gz</tt> and
 <tt>MEI_Callset_T2T-CHM13.ALL.20241211.csv.gz</tt>) and converted to
 bigBed following the steps described in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/mei.txt"
 target="_blank">makeDoc file</a>. Conversion uses scripts in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/mei"
-target="_blank">src/hg/makeDb/scripts/mei</a>: VCF-style positions
+target="_blank">src/hg/makeDb/scripts/mei</a>,
+and the track configuration is in <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/mei.ra" target="_blank">trackDb/human/mei.ra</a>: VCF-style positions
 (1-based POS, anchor base) are converted to half-open BED coordinates
 (<tt>chromStart = POS - 1</tt>, <tt>chromEnd = chromStart + 1</tt>),
 genotypes are tallied across the 65 samples, and items are colored by
 mobile element class.
 </p>
 
 <h2>Data Access</h2>
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from
 there to spreadsheet or tab-separated tables. From scripts, the data can
-be accessed through our <a href="https://api.genome.ucsc.edu">API</a>,
+be accessed through our <a href="https://api.genome.ucsc.edu" target="_blank">API</a>,
 track=<i>meiHgsvc3</i>.
 </p>
 <p>
 For automated download and analysis, the genome annotation is stored in
 a bigBed file that can be downloaded from
-<a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/" target="_blank">
+<a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/" target="_blank">
 our download server</a>.  The file for this track is called
-<tt>hgsvc3.bb</tt> in <tt>/gbdb/hg38/mei/</tt> (GRCh38) or
-<tt>/gbdb/hs1/mei/</tt> (T2T-CHM13). Individual regions or the whole
+<tt>hgsvc3.bb</tt> in <tt>/gbdb/$db/mei/</tt>. Individual regions or the whole
 genome annotation can be obtained using our tool <tt>bigBedToBed</tt>,
 which can be compiled from the source code or downloaded as a
 precompiled binary for your system. Instructions for downloading source
 code and binaries can be found
-<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">here</a>.
+<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" target="_blank">here</a>.
 The tool can also be used to obtain features within a given range, e.g.
-<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/hgsvc3.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>.
+<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/hgsvc3.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>.
 </p>
 <p>
 The original annotation source data can be downloaded from the
 <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/HGSVC3/release/Mobile_Elements/1.0/"
 target="_blank">HGSVC3 1000 Genomes FTP site</a>.
 </p>
 
 <h2>Credits</h2>
 <p>
 Thanks to the Human Genome Structural Variation Consortium phase 3
 (HGSVC3) for releasing the underlying assemblies and MEI callsets used
 to produce this track.
 </p>
 
 <h2>References</h2>