7e87cadb469b4e0eb4fb7f973154357cfe7fc345 lrnassar Mon Sep 21 15:56:18 2026 -0700 QA fixes for the mei (Mobile Insertions) track collection. refs #37524 Fix two data bugs found during QA and rebuild the affected bigBeds. meiEul1dbToBed.py looked up samples and individuals by name, but euL1db joins on 1-based row numbers, so neither join ever matched and the individual count, tissues, clinical conditions and populations were empty on all 8,991 insertions while the contributing-samples table printed row numbers. Both loaders now key on the row number, the table prints the sample name, and the adjacent population filter is case-insensitive so it actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every record, which dropped the first base of the element on the 96 GRCh38 and 111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no anchor base; it now prefers INFO SEQ, which always matches SVLEN. Correct seven statements on the description pages against their sources: the HGSVC3 single-caller split was attributed to PALMER rather than L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db was credited with aligning the L1HS consensus when the paper says it was processed from our RepeatMasker track, DeepMEI's network was described as a classifier rather than a genotyper and given the wrong training set, euL1db listed two detection methods absent from the data, and HMEID contradicted itself on the MELT ASSESS cutoff. Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted callset is kept off the download server; the container page no longer claims the whole collection is long-read, lists the two euL1db subtracks, scopes its display conventions to the subtracks they describe, and cites all six papers; dead and wrong track links are repointed and pinned to a db; $db replaces hardcoded hg38 in paths on pages that serve three assemblies; the euL1db labels no longer carry hg38 counts and a lift note that made no sense on hg19; all six subtracks gain a dataVersion; the euL1db filter ranges match the data; and five autoSql field descriptions match what the files contain. Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19 carries the two euL1db subtracks. diff --git src/hg/makeDb/trackDb/human/meiHgsvc3.html src/hg/makeDb/trackDb/human/meiHgsvc3.html index e1711b8e13b..94230b9fc1d 100644 --- src/hg/makeDb/trackDb/human/meiHgsvc3.html +++ src/hg/makeDb/trackDb/human/meiHgsvc3.html @@ -33,31 +33,31 @@ of the inserted mobile element (the ALT allele minus the anchor base). </p> <h2>Display Conventions and Configuration</h2> <p> An insertion has zero length on the reference: it attaches between two adjacent reference bases without replacing any of them. Following the VCF convention used by the underlying PAV calls and by the other long-read SV tracks, each MEI is drawn as a <b>1-bp block sitting on the anchor base</b> — the reference base immediately to the left of the insertion attachment point. The inserted mobile element itself is not present in the reference and is therefore not drawn; its length is reported on the detail page (svLen) and in the item label (<tt>class-svLen:carrierCount</tt>). bigBed does support truly zero-width features between nucleotides, but for consistency with the -<a href="hgTrackUi?g=longReadVariants">long-read SV tracks</a>, this track uses the +<a href="hgTrackUi?db=hg38&g=longReadVariants">long-read SV tracks</a>, this track uses the 1-bp anchor representation instead. </p> <p> Items are colored by element class: </p> <ul> <li><span style="display:inline-block;background-color:#0072B2;width:18px;height:12px;vertical-align:middle;"></span> <b>Alu</b> — SINE (Short INterspersed Element)</li> <li><span style="display:inline-block;background-color:#D55E00;width:18px;height:12px;vertical-align:middle;"></span> <b>L1</b> — LINE-1 (Long INterspersed Element-1)</li> <li><span style="display:inline-block;background-color:#009E73;width:18px;height:12px;vertical-align:middle;"></span> <b>SVA</b> (SINE-VNTR-Alu) — composite retrotransposon</li> <li><span style="display:inline-block;background-color:#CC79A7;width:18px;height:12px;vertical-align:middle;"></span> <b>HERVK</b> (Human Endogenous Retrovirus K) — endogenous retrovirus</li> <li><span style="display:inline-block;background-color:#000000;width:18px;height:12px;vertical-align:middle;"></span> <b>snRNA</b> — small nuclear RNA</li> </ul> <p> The score column encodes the alt-allele frequency on a 0-1000 scale. Filters allow restricting to specific element classes, length ranges, @@ -65,85 +65,86 @@ and reference repeat overlap. </p> <h2>Methods</h2> <p> The HGSVC3 study sequenced and de novo assembled 65 individuals (30 males, 35 females) representing five continental groups and 28 populations: 30 of African, 9 of Admixed American, 8 of European, 10 of East Asian and 8 of South Asian descent, with three parent-child trios included. Each sample was sequenced to ~47-fold coverage of PacBio HiFi and ~56-fold coverage of Oxford Nanopore long reads (~36-fold ultra-long), and complemented with Strand-seq, Bionano optical mapping, Hi-C, Iso-Seq and RNA-seq. Haplotype-resolved diploid assemblies were produced and structural variants called with PAV. Mobile element insertions were identified from the union of two independent MEI -callsets, L1ME-AID and PALMER; all single-caller calls were manually +callsets, L1ME-AID and PALMER2; all single-caller calls were manually curated. Orthogonal validation against an independent MELT-LRA callset -showed an average concordance of 90.8% on GRCh38. Roughly 93% of MEIs -are supported by both callers; the remaining single-caller calls split -about 6:1 in favour of PALMER (PALMER-only ~6%, L1ME-AID-only ~1%). +showed a concordance of 92.5% on GRCh38 and 92.1% on T2T-CHM13. Roughly 93% of MEIs +are supported by both callers (93.6% on GRCh38, 92.4% on T2T-CHM13); the +remaining single-caller calls split about 6:1 in favour of L1ME-AID +(L1ME-AID-only 5.5% on GRCh38, PALMER-only 0.9%). The Caller Count, PALMER Validated and L1ME-AID Validated filters can be used to restrict the display to the high-confidence dual-validated subset. Calls are restricted to non-low-confidence regions (i.e. excluding Yq12 and centromeres). For each site, per-sample genotypes from all 65 assembled samples are summarized into an alt-allele count, allele number, allele frequency and a list of carrier samples. See Logsdon et al. 2025 (Nature) for full methodological details. </p> <p> The original CSVs were downloaded from the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/HGSVC3/release/Mobile_Elements/1.0/" target="_blank">HGSVC3 Mobile Elements release directory</a> (files <tt>MEI_Callset_GRCh38.ALL.20241211.csv.gz</tt> and <tt>MEI_Callset_T2T-CHM13.ALL.20241211.csv.gz</tt>) and converted to bigBed following the steps described in the <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/mei.txt" target="_blank">makeDoc file</a>. Conversion uses scripts in <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/mei" -target="_blank">src/hg/makeDb/scripts/mei</a>: VCF-style positions +target="_blank">src/hg/makeDb/scripts/mei</a>, +and the track configuration is in <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/mei.ra" target="_blank">trackDb/human/mei.ra</a>: VCF-style positions (1-based POS, anchor base) are converted to half-open BED coordinates (<tt>chromStart = POS - 1</tt>, <tt>chromEnd = chromStart + 1</tt>), genotypes are tallied across the 65 samples, and items are colored by mobile element class. </p> <h2>Data Access</h2> <p> The data can be explored interactively in table format with the <a href="../cgi-bin/hgTables">Table Browser</a> or the <a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to spreadsheet or tab-separated tables. From scripts, the data can -be accessed through our <a href="https://api.genome.ucsc.edu">API</a>, +be accessed through our <a href="https://api.genome.ucsc.edu" target="_blank">API</a>, track=<i>meiHgsvc3</i>. </p> <p> For automated download and analysis, the genome annotation is stored in a bigBed file that can be downloaded from -<a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/" target="_blank"> +<a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/" target="_blank"> our download server</a>. The file for this track is called -<tt>hgsvc3.bb</tt> in <tt>/gbdb/hg38/mei/</tt> (GRCh38) or -<tt>/gbdb/hs1/mei/</tt> (T2T-CHM13). Individual regions or the whole +<tt>hgsvc3.bb</tt> in <tt>/gbdb/$db/mei/</tt>. Individual regions or the whole genome annotation can be obtained using our tool <tt>bigBedToBed</tt>, which can be compiled from the source code or downloaded as a precompiled binary for your system. Instructions for downloading source code and binaries can be found -<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">here</a>. +<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads" target="_blank">here</a>. The tool can also be used to obtain features within a given range, e.g. -<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/mei/hgsvc3.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>. +<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/$db/mei/hgsvc3.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>. </p> <p> The original annotation source data can be downloaded from the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/HGSVC3/release/Mobile_Elements/1.0/" target="_blank">HGSVC3 1000 Genomes FTP site</a>. </p> <h2>Credits</h2> <p> Thanks to the Human Genome Structural Variation Consortium phase 3 (HGSVC3) for releasing the underlying assemblies and MEI callsets used to produce this track. </p> <h2>References</h2>