7e87cadb469b4e0eb4fb7f973154357cfe7fc345
lrnassar
  Mon Sep 21 15:56:18 2026 -0700
QA fixes for the mei (Mobile Insertions) track collection. refs #37524

Fix two data bugs found during QA and rebuild the affected bigBeds.
meiEul1dbToBed.py looked up samples and individuals by name, but euL1db
joins on 1-based row numbers, so neither join ever matched and the
individual count, tissues, clinical conditions and populations were empty
on all 8,991 insertions while the contributing-samples table printed row
numbers. Both loaders now key on the row number, the table prints the
sample name, and the adjacent population filter is case-insensitive so it
actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every
record, which dropped the first base of the element on the 96 GRCh38 and
111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no
anchor base; it now prefers INFO SEQ, which always matches SVLEN.

Correct seven statements on the description pages against their sources:
the HGSVC3 single-caller split was attributed to PALMER rather than
L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db
was credited with aligning the L1HS consensus when the paper says it was
processed from our RepeatMasker track, DeepMEI's network was described as
a classifier rather than a genotyper and given the wrong training set,
euL1db listed two detection methods absent from the data, and HMEID
contradicted itself on the MELT ASSESS cutoff.

Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted
callset is kept off the download server; the container page no longer
claims the whole collection is long-read, lists the two euL1db subtracks,
scopes its display conventions to the subtracks they describe, and cites
all six papers; dead and wrong track links are repointed and pinned to a
db; $db replaces hardcoded hg38 in paths on pages that serve three
assemblies; the euL1db labels no longer carry hg38 counts and a lift note
that made no sense on hg19; all six subtracks gain a dataVersion; the
euL1db filter ranges match the data; and five autoSql field descriptions
match what the files contain.

Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct
the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19
carries the two euL1db subtracks.

diff --git src/hg/makeDb/trackDb/human/meiDeepmei1kg.html src/hg/makeDb/trackDb/human/meiDeepmei1kg.html
index eb4171495cc..9ba75c0e97a 100644
--- src/hg/makeDb/trackDb/human/meiDeepmei1kg.html
+++ src/hg/makeDb/trackDb/human/meiDeepmei1kg.html
@@ -1,26 +1,27 @@
 <h2>Description</h2>
 <p>
 This track shows <b>mobile element insertions (MEIs)</b> called by
 <a href="https://github.com/xuxif/DeepMEI" target="_blank">DeepMEI</a>
 on the 3,202 high-coverage 1000 Genomes Project samples (NYGC
 re-sequencing) aligned to GRCh38. At each site, at least one of the
 3,202 samples carries a non-reference insertion of an Alu, L1 (LINE-1)
-or SVA mobile element. DeepMEI is a convolutional neural-network
-caller that scans short-read alignments for the read-pair, split-read
-and clipping signatures of a new insertion and classifies each
-candidate site as Alu, L1 or SVA.
+or SVA mobile element. DeepMEI screens short-read alignments for
+the soft-clipped and discordant read signatures of a new insertion,
+assigns each candidate to Alu, L1 or SVA by aligning it to mobile
+element consensus sequences, and then genotypes the site with a
+convolutional neural network.
 </p>
 
 <table class="stdTbl">
 <tr><th>Class</th><th>MEIs</th></tr>
 <tr><td>Alu</td><td>68,282</td></tr>
 <tr><td>L1</td><td>16,891</td></tr>
 <tr><td>SVA</td><td>6,444</td></tr>
 <tr><th>Total</th><th>91,617</th></tr>
 </table>
 
 <p>
 For each MEI, the track lists the element class, the alt-allele count,
 allele number and allele frequency across the 3,202 samples, the number
 of carrier samples, and the list of carrier sample IDs.
 </p>
@@ -39,85 +40,99 @@
 on this track. The item label is <tt>class-carrierCount</tt>.
 </p>
 <p>
 Items are colored by element class:
 </p>
 <ul>
   <li><span style="display:inline-block;background-color:#0072B2;width:18px;height:12px;vertical-align:middle;"></span> <b>Alu</b> &mdash; SINE (Short INterspersed Element)</li>
   <li><span style="display:inline-block;background-color:#D55E00;width:18px;height:12px;vertical-align:middle;"></span> <b>L1</b> &mdash; LINE-1 (Long INterspersed Element-1)</li>
   <li><span style="display:inline-block;background-color:#009E73;width:18px;height:12px;vertical-align:middle;"></span> <b>SVA</b> (SINE-VNTR-Alu) &mdash; composite retrotransposon</li>
 </ul>
 <p>
 The score column encodes the alt-allele frequency on a 0-1000 scale.
 Filters allow restricting to specific element classes, allele frequency
 and carrier counts.
 </p>
+<p>
+This callset carries no site-level quality filter: the FILTER column is
+empty on all 91,617 records. It is also singleton-heavy, with 43,627
+sites (47.6%) called in a single sample. Raising the <b>Carrier Sample
+Count</b> filter above 1 restricts the display to sites seen in more
+than one individual.
+</p>
 
 <h2>Methods</h2>
 <p>
-DeepMEI is a deep convolutional neural network that detects
-non-reference mobile element insertions from short-read whole-genome
-sequencing. For every candidate site supported by an anomalous
-read-pair, split-read or soft-clip signature, the surrounding alignment
-pile-up is encoded as an image and passed through a CNN that classifies
-the site as Alu, L1, SVA or background. The model was trained on
-labelled MEIs from the 1000 Genomes phase 3 callset and orthogonal
-long-read truth sets. For this track, DeepMEI was run on the
+DeepMEI detects non-reference mobile element insertions from
+short-read whole-genome sequencing. Candidate sites are found from
+soft-clipped and discordant reads, and each candidate is assigned to
+Alu, L1 or SVA by aligning the soft-clipped sequence against a
+library of mobile element consensus sequences from Repbase. The surrounding alignment pile-up is then encoded
+as a six-channel 351 &times; 100 image and passed through an Inception
+v3 convolutional neural network, which returns probabilities for the
+three genotypes homozygous reference, heterozygous and homozygous
+insertion; the highest probability becomes the genotype call. The model
+was trained on ten samples from the 1000 Genomes Project phase 4
+release, using the project's freeze_V3 structural variant calls as
+labels after manual confirmation in IGV. Long-read data was used to
+benchmark the tool rather than to train it. For this track, DeepMEI was
+run on the
 high-coverage (~30&times;) Illumina re-sequencing of all 3,202
 1000 Genomes Project samples produced by the New York Genome Center
 (NYGC), giving 6,404 haplotypes per site. See Xu et al. 2023 (bioRxiv)
 for full methodological details.
 </p>
 
 <p>
 The original VCF was downloaded from the DeepMEI GitHub repository
 (file <tt>merge_1000g.latested.vcf.gz</tt> in
 <a href="https://github.com/xuxif/DeepMEI/tree/main/DeepMEI/1000g_high_callset"
 target="_blank">DeepMEI/1000g_high_callset/</a>) and converted to
 bigBed following the steps described in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/mei.txt"
 target="_blank">makeDoc file</a>. Conversion uses scripts in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/mei"
-target="_blank">src/hg/makeDb/scripts/mei</a>: VCF-style positions
+target="_blank">src/hg/makeDb/scripts/mei</a>,
+and the track configuration is in <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/mei.ra" target="_blank">trackDb/human/mei.ra</a>: VCF-style positions
 (1-based POS, anchor base) are converted to half-open BED coordinates
 (<tt>chromStart = POS - 1</tt>, <tt>chromEnd = chromStart + 1</tt>),
 per-sample genotypes are tallied across the 3,202 samples, and items
 are colored by mobile element class.
 </p>
 
 <h2>Data Access</h2>
 <p>
 The data can be explored interactively in table format with the
 <a href="../cgi-bin/hgTables">Table Browser</a> or the
 <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from
 there to spreadsheet or tab-separated tables. From scripts, the data can
-be accessed through our <a href="https://api.genome.ucsc.edu">API</a>,
+be accessed through our <a href="https://api.genome.ucsc.edu" target="_blank">API</a>,
 track=<i>meiDeepmei1kg</i>.
 </p>
 <p>
 For automated download and analysis, the genome annotation is stored in
 a bigBed file that can be downloaded from
-<a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/" target="_blank">
+<a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/" target="_blank">
 our download server</a>.  The file for this track is called
-<tt>deepmei1kg.bb</tt> in <tt>/gbdb/hg38/mei/</tt>. Individual regions
+<tt>deepmei1kg.bb</tt> in <tt>/gbdb/$db/mei/</tt>. Individual regions
 or the whole genome annotation can be obtained using our tool
 <tt>bigBedToBed</tt>, which can be compiled from the source code or
 downloaded as a precompiled binary for your system. Instructions for
 downloading source code and binaries can be found
-<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">here</a>.
+<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" target="_blank">here</a>.
 The tool can also be used to obtain features within a given range, e.g.
-<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/deepmei1kg.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>.
+<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/deepmei1kg.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>.
 </p>
 <p>
 The original annotation source data can be downloaded from the
 <a href="https://github.com/xuxif/DeepMEI/tree/main/DeepMEI/1000g_high_callset"
 target="_blank">DeepMEI GitHub repository</a>.
 </p>
 
 <h2>Credits</h2>
 <p>
 Thanks to Xiaofei Xu, Fengxiao Bu and colleagues for developing
 DeepMEI and releasing the 1000 Genomes MEI callset, and to the New
 York Genome Center for producing the underlying high-coverage
 1000 Genomes re-sequencing data.
 </p>