7e87cadb469b4e0eb4fb7f973154357cfe7fc345 lrnassar Mon Sep 21 15:56:18 2026 -0700 QA fixes for the mei (Mobile Insertions) track collection. refs #37524 Fix two data bugs found during QA and rebuild the affected bigBeds. meiEul1dbToBed.py looked up samples and individuals by name, but euL1db joins on 1-based row numbers, so neither join ever matched and the individual count, tissues, clinical conditions and populations were empty on all 8,991 insertions while the contributing-samples table printed row numbers. Both loaders now key on the row number, the table prints the sample name, and the adjacent population filter is case-insensitive so it actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every record, which dropped the first base of the element on the 96 GRCh38 and 111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no anchor base; it now prefers INFO SEQ, which always matches SVLEN. Correct seven statements on the description pages against their sources: the HGSVC3 single-caller split was attributed to PALMER rather than L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db was credited with aligning the L1HS consensus when the paper says it was processed from our RepeatMasker track, DeepMEI's network was described as a classifier rather than a genotyper and given the wrong training set, euL1db listed two detection methods absent from the data, and HMEID contradicted itself on the MELT ASSESS cutoff. Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted callset is kept off the download server; the container page no longer claims the whole collection is long-read, lists the two euL1db subtracks, scopes its display conventions to the subtracks they describe, and cites all six papers; dead and wrong track links are repointed and pinned to a db; $db replaces hardcoded hg38 in paths on pages that serve three assemblies; the euL1db labels no longer carry hg38 counts and a lift note that made no sense on hg19; all six subtracks gain a dataVersion; the euL1db filter ranges match the data; and five autoSql field descriptions match what the files contain. Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19 carries the two euL1db subtracks. diff --git src/hg/makeDb/trackDb/human/meiEul1dbRef.html src/hg/makeDb/trackDb/human/meiEul1dbRef.html index af70e9a4b85..d575647338a 100644 --- src/hg/makeDb/trackDb/human/meiEul1dbRef.html +++ src/hg/makeDb/trackDb/human/meiEul1dbRef.html @@ -1,94 +1,99 @@

Description

This track shows the L1-HS (human-specific LINE-1) retrotransposon copies that are already present in the human reference genome, as catalogued by euL1db. Unlike the companion -euL1db Insertions track, which shows polymorphic +euL1db Insertions track, which shows polymorphic insertions not present (or variably present) in the reference, the items here are the L1-HS copies that the reference genome already contains. Together, the two tracks provide a comprehensive view of L1-HS positions in the genome — both the "fixed" reference set and the variants segregating in the human population.

Display Conventions and Configuration

Each item is a single reference L1-HS element. The name shows the L1HS sub-group. Items are coloured by sub-group:

L1HS-Ta — youngest, currently active L1-HS subset
L1HS-PreTa — older L1-HS subset, ancestral to Ta
L1HS-undef — sub-group not assigned in euL1db

The detail page for each element shows the element length, its integrity (full-length, 5′-truncated, 3′-truncated, internal_fragment), and its position on the L1HS consensus sequence. Filters are provided for sub-group and integrity.

Methods

-The reference L1-HS catalogue was assembled by the euL1db curators by aligning -the L1HS consensus sequence to the hg19 reference genome and classifying each -match into one of the L1HS sub-groups. See Mir et al. 2015 for further details. +The reference L1-HS catalogue was assembled by the euL1db curators from the +UCSC RepeatMasker track table for hg19: the L1HS +annotations were extracted and each element was annotated with its sub-group, +integrity and position on the L1HS consensus sequence. euL1db uses this table +internally to decide whether a given insertion polymorphism corresponds to an +L1HS copy already present in the reference. See Mir et al. 2015 for further +details.

Track files were generated from the euL1db v1.00 ReferenceL1HS.txt table (data dump downloaded March 2018) using the script meiEul1dbRefToBed.py. The hg19 BED was lifted to hg38 with liftOver; of 1,544 hg19 elements, 1,540 lifted successfully. For details see -hg19/mei.txt and the scripts directory +doc/hg38/mei.txt and the scripts directory src/hg/makeDb/scripts/mei. + target="_blank">src/hg/makeDb/scripts/mei, + and the track configuration is in trackDb/human/mei.ra.

Data Access

The data can be explored interactively in table format with the Table Browser or the Data Integrator and exported from there to spreadsheet or tab-separated tables. From scripts, the data can be accessed through our REST API, track=meiEul1dbRef.

For automated download and analysis, the annotation is stored in a bigBed file that can be downloaded from -our download +our download server. The file for this track is called eul1dbRef.bb. Individual regions or the whole genome annotation can be obtained using our tool bigBedToBed.

The original annotation source data can be downloaded from eul1db.unice.fr via the Download tab.

Credits

Thanks to Gaël Cristofari and colleagues at IRCAN (Nice, France) for making the euL1db data freely available. Track built at UCSC by the Genome Browser group.

References

Mir AA, Philippe C, Cristofari G. euL1db: the European database of L1HS retrotransposon insertions in humans. Nucleic Acids Res. 2015 Jan;43(Database issue):D43-7. PMID: 25352549; PMC: PMC4383891