7e87cadb469b4e0eb4fb7f973154357cfe7fc345 lrnassar Mon Sep 21 15:56:18 2026 -0700 QA fixes for the mei (Mobile Insertions) track collection. refs #37524 Fix two data bugs found during QA and rebuild the affected bigBeds. meiEul1dbToBed.py looked up samples and individuals by name, but euL1db joins on 1-based row numbers, so neither join ever matched and the individual count, tissues, clinical conditions and populations were empty on all 8,991 insertions while the contributing-samples table printed row numbers. Both loaders now key on the row number, the table prints the sample name, and the adjacent population filter is case-insensitive so it actually drops "unknown". meiHgsvc3CsvToBed.py took alt[1:] on every record, which dropped the first base of the element on the 96 GRCh38 and 111 T2T-CHM13 records where PALMER2 is the only caller and ALT carries no anchor base; it now prefers INFO SEQ, which always matches SVLEN. Correct seven statements on the description pages against their sources: the HGSVC3 single-caller split was attributed to PALMER rather than L1ME-AID, its orthogonal concordance was 90.8% rather than 92.5%, euL1db was credited with aligning the L1HS consensus when the paper says it was processed from our RepeatMasker track, DeepMEI's network was described as a classifier rather than a genotyper and given the wrong training set, euL1db listed two detection methods absent from the data, and HMEID contradicted itself on the MELT ASSESS cutoff. Also: the SweGen bigDataUrl now points at _swegen.bb so the restricted callset is kept off the download server; the container page no longer claims the whole collection is long-read, lists the two euL1db subtracks, scopes its display conventions to the subtracks they describe, and cites all six papers; dead and wrong track links are repointed and pinned to a db; $db replaces hardcoded hg38 in paths on pages that serve three assemblies; the euL1db labels no longer carry hg38 counts and a lift note that made no sense on hg19; all six subtracks gain a dataVersion; the euL1db filter ranges match the data; and five autoSql field descriptions match what the files contain. Document the gbdb symlinks and the QA changes in doc/hg38/mei.txt, correct the HMEID bedToBigBed type there, and add an hg19.txt pointer since hg19 carries the two euL1db subtracks. diff --git src/hg/makeDb/doc/hg38/mei.txt src/hg/makeDb/doc/hg38/mei.txt index 14a96541bec..51ad3ef85d1 100644 --- src/hg/makeDb/doc/hg38/mei.txt +++ src/hg/makeDb/doc/hg38/mei.txt @@ -120,31 +120,31 @@ wget http://bigdata.ibp.ac.cn/HMEID/static/download/sample_info.HMEIDv1.1.txt.gz # Convert site-level VCF (36699 MEIs) to bed9+27. # Source: ~/kent/src/hg/makeDb/scripts/mei/meiHmeidVcfToBed.py python3 ~/kent/src/hg/makeDb/scripts/mei/meiHmeidVcfToBed.py \ MEI.GRCh38.HMEIDv1.1.vcf.gz \ /hive/data/genomes/hg38/chrom.sizes \ meiHmeid.bed # -> Read 36699 records, wrote 36699, skipped 0 + 0 + 0 # Class distribution: Alu 26553, HERVK 126, L1 7353, SVA 2667. sort -k1,1 -k2,2n meiHmeid.bed > meiHmeid.sorted.bed bedToBigBed -tab \ -as=$HOME/kent/src/hg/makeDb/scripts/mei/meiHmeid.as \ - -type=bed9+27 \ + -type=bed9+28 \ meiHmeid.sorted.bed \ /hive/data/genomes/hg38/chrom.sizes \ meiHmeid.bb ############################################################ # 2026-05-13 Claude (max) - SweGen MELT MEI callset (hg38, lifted from hg19) # Source: Ameur et al. 2017, Eur J Hum Genet, PMID 28832569 (SweGen cohort) # Gardner et al. 2017, Genome Res, PMID 28855259 (MELT tool) # https://swefreq.nbis.se/dataset/SweGen # Site-level MELT v2.0.2 MEI callset on 1,000 Swedish WGS samples # (SweGen, Illumina HiSeq X, 150 bp PE, BWA-MEM v0.7.12 to GRCh37). # The VCF has no per-sample columns; INFO carries MELT_AN (allele # count, despite the name) and MELT_AF (allele frequency) plus the # usual MELT fields SVTYPE/SVLEN/TSD/ASSESS/MEIINFO/INTERNAL. @@ -260,15 +260,74 @@ bedToBigBed -tab \ -as=$HOME/kent/src/hg/makeDb/scripts/mei/meiEul1db.as \ -type=bed9+19 \ /hive/data/genomes/hg38/bed/mei/eul1db/eul1db.hg38.sorted.bed \ /hive/data/genomes/hg38/chrom.sizes \ /hive/data/genomes/hg38/bed/mei/eul1db/eul1db.hg38.bb sort -k1,1 -k2,2n eul1dbRef.hg38.bed \ > /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.sorted.bed bedToBigBed -tab \ -as=$HOME/kent/src/hg/makeDb/scripts/mei/meiEul1dbRef.as \ -type=bed9+6 \ /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.sorted.bed \ /hive/data/genomes/hg38/chrom.sizes \ /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.bb + +############################################################ +# /gbdb symlinks for the whole collection +# The trackDb stanza lives in trackDb/human/mei.ra and uses $D, so each +# assembly needs its own symlink directory. The SweGen file is +# underscore-prefixed because the callset cannot be redistributed (see +# below), which keeps it off the public download server; the stanza also +# carries "tableBrowser off", which blocks both hgTables and the REST API. + +mkdir -p /gbdb/hg38/mei /gbdb/hs1/mei /gbdb/hg19/mei + +ln -s /hive/data/genomes/hg38/bed/mei/meiHgsvc3.bb /gbdb/hg38/mei/hgsvc3.bb +ln -s /hive/data/genomes/hg38/bed/mei/deepmei/deepmei1kg.bb /gbdb/hg38/mei/deepmei1kg.bb +ln -s /hive/data/genomes/hg38/bed/mei/hmei/meiHmeid.bb /gbdb/hg38/mei/hmeid.bb +ln -s /hive/data/genomes/hg38/bed/mei/swegen/meiSwegen.bb /gbdb/hg38/mei/_swegen.bb +ln -s /hive/data/genomes/hg38/bed/mei/eul1db/eul1db.hg38.bb /gbdb/hg38/mei/eul1db.bb +ln -s /hive/data/genomes/hg38/bed/mei/eul1db/eul1dbRef.hg38.bb /gbdb/hg38/mei/eul1dbRef.bb +ln -s /hive/data/genomes/hs1/bed/mei/meiHgsvc3.bb /gbdb/hs1/mei/hgsvc3.bb +ln -s /hive/data/genomes/hg19/bed/mei/eul1db/eul1db.hg19.bb /gbdb/hg19/mei/eul1db.bb +ln -s /hive/data/genomes/hg19/bed/mei/eul1db/eul1dbRef.hg19.bb /gbdb/hg19/mei/eul1dbRef.bb + +############################################################ +# 2026-09-21 Lou - QA fixes (refs #37524) + +# 1) meiHgsvc3: inserted sequence on the PALMER-only records. +# The source CSV merges three different INFO schemas. Most records follow +# the VCF convention (ALT[0] == REF[0], len(ALT)-1 == SVLEN), but the 96 +# GRCh38 / 111 T2T-CHM13 records whose INFO carries SEQ/CALLERS do not: +# there ALT is the element itself (len(ALT) == SVLEN, and on 70 of the 96 +# ALT[0] differs from REF[0]), and 13 of them carry a truncated ALT. +# meiHgsvc3CsvToBed.py took alt[1:] unconditionally, which dropped the +# first base of the element on all 96 and left insertSeq disagreeing with +# svLen. The script now prefers INFO SEQ when present, which always +# matches SVLEN. Rebuilt both assemblies with the commands above; +# item counts and every other column are unchanged. + +# 2) meiEul1db: sample and individual joins. +# euL1db's foreign keys are 1-based row numbers, not names: SRIP.txt's +# sampleID column indexes Samples.txt (943 rows) and Samples.txt's +# Individual_id column indexes Individuals.txt (741 rows). Both totals +# match Table 1 of Mir et al. 2015. meiEul1dbToBed.py keyed both lookup +# tables by name, so neither join ever matched and individualCount, +# tissues, diseases and populations were empty on all 8,991 MRIPs, while +# the contributing-samples table printed row numbers instead of sample +# names. Both loaders now key on the row number and the sample table +# prints the name. The adjacent population filter compared against +# "Unknown" while the data says "unknown", so it never fired either; it is +# now case-insensitive. After the fix: individualCount set on 8,976 +# records, tissues 8,976, diseases 6,781, populations 7,589 (hg19; 7,586 +# after the lift to hg38, which loses 3 MRIPs). Rebuilt the +# hg19 bigBed and re-lifted to hg38 with the commands above; item counts +# (8,991 / 8,988) and all other columns are unchanged. + +# 3) meiSwegen: /gbdb/hg38/mei/swegen.bb renamed to _swegen.bb. The +# callset cannot be redistributed, and the underscore prefix is what keeps +# a file off the public download server. "tableBrowser off" already +# blocked hgTables and the REST API (verified: the API returns HTTP 403 +# "protected data"), but without the prefix the bigBed itself would have +# been mirrored to hgdownload. mei.ra was updated to match.