0bd565e053abc8c74475f352bedfd37e41312fd2 max Wed Aug 12 02:15:40 2026 -0700 lrSv: update noyvertSv docs and align merged-track source labels, refs #37888 Follow-up to author (Boris Noyvert) feedback on the Noyvert/Boehringer long-read SV dataset. noyvertSv.html: - restore neutral wording about the shared 1000G ONT reads; drop the "independent reprocessing" phrasing and the call-level overlap interpretation the authors objected to - note that singletons (SVs in a single sample) were excluded, so the panel is not exhaustive for the rarest variants - add the medRxiv preprint link alongside the eLife reference Give each dataset one consistent name across its subtrack and the merged (lrSvAll) source filter (databases.tsv + lrSvAll.ra + lrSv.ra): Noyvert 888 (1000G ONT) -> 1KG ONT Boehringer 888 1KG ONT Vienna 1,019 -> 1KG ONT 1019 1KG ONT 100 (Gustafson) -> 1KG ONT UW 100 The gustafsonSv subtrack short/long labels read 97 samples; the paper and our track docs report 100 (Gustafson et al. 2024, PMID 39358015), so those are corrected to 100 as well. Rebuilt lrSvAll.bb with lrSvMergeAll.py; item count unchanged (2,582,278). diff --git src/hg/makeDb/trackDb/human/noyvertSv.html src/hg/makeDb/trackDb/human/noyvertSv.html index 488b8d9039e..e4547b64013 100644 --- src/hg/makeDb/trackDb/human/noyvertSv.html +++ src/hg/makeDb/trackDb/human/noyvertSv.html @@ -5,50 +5,44 @@ Genomes Project, representing five ancestry groups. This dataset and the 1KG ONT Vienna track (Schloissnig et al. 2025) are based on the same underlying Oxford Nanopore sequencing data; the 888 samples here are a subset of the 1,019 samples in that track and only SVs that appear in a single sample (singletons) were removed from this track, so this callset is smaller than the Schloissnig dataset. The reason is that this callset was created primarily for imputation: The SVs here were merged with previously identified short variants from the same individuals to generate a multi-ancestry SV imputation reference panel. This panel was used to impute SVs in approximately 500,000 UK Biobank participants and test their associations with 32 disease-relevant traits.

The track contains all 107,445 SVs in the reference panel: 59,953 insertions, 38,459 deletions, 5,729 inversions, 2,696 breakends, and 608 duplications. -Variants seen in only a single individual (singletons) were excluded from the -panel, so every SV shown was observed in at least two individuals; the panel is -therefore not exhaustive for very rare variants. Each variant is annotated with its overall allele frequency; allele frequencies across five superpopulations (African, Admixed American, East Asian, European, and South Asian); Hardy-Weinberg equilibrium p-values; and imputation accuracy metrics from internal leave-one-out validation and UK Biobank imputation. For SVs reaching genome-wide significance, the associated traits, p-values, and INFO scores are listed on the corresponding variant details page.

-This dataset and the 1KG ONT Vienna track -(Schloissnig et al. 2025) are based on the same underlying Oxford Nanopore -sequencing data; the 888 samples here are a subset of the 1,019 samples in that -track. The two studies applied different data-processing and SV-calling -pipelines to address distinct research objectives, so the individual calls are -only partially concordant. The imputation reference panel, UK Biobank -imputation results, and SV-wide association study (SV-WAS) results described -here are specific to this track. +Although the two studies share the same raw sequencing data, they applied +different data-processing and SV-calling pipelines to address distinct research +objectives, so the individual calls are only partially concordant. The +imputation reference panel, UK Biobank imputation results, and SV-wide +association study (SV-WAS) results described here are specific to this track.

Display Conventions and Configuration

Items are colored by SV type, matching the other subtracks of the container:

  Deletion (DEL)
  Insertion (INS)
  Duplication (DUP)
  Inversion (INV)