0656b0a9ad98986ef3944b6ceaf305ab08ca4e39 max Fri Jul 24 01:10:13 2026 -0700 noyvertSv: swap UK Biobank r2 filter for leave-one-out metrics, revise description Per author (Boris Noyvert) request, replace the r2Ukb (UK Biobank imputation r2) filter with r2Loo and concordanceLoo, the primary SV imputation quality measures for the multi-ancestry reference panel (r2Ukb mainly reflects European-ancestry performance). Both fields already exist in noyvert.bb, so this is a trackDb-only change. Also update the description page with the author's revised text, including a comparison with the 1KG ONT Vienna dataset. refs #37888 diff --git src/hg/makeDb/trackDb/human/noyvertSv.html src/hg/makeDb/trackDb/human/noyvertSv.html index acb0734d582..d37b877d001 100644 --- src/hg/makeDb/trackDb/human/noyvertSv.html +++ src/hg/makeDb/trackDb/human/noyvertSv.html @@ -1,37 +1,44 @@

Description

-This track shows structural variants (SVs) identified by Oxford Nanopore -long-read sequencing of 888 individuals from the 1000 Genomes Project, -spanning five ancestry groups. Structural variants are genomic rearrangements -larger than about 50 bp, such as deletions, insertions, inversions, -duplications and breakends (rearrangement junctions); because they alter or -move large stretches of DNA at once, they can affect gene dosage and gene -regulation more strongly than single-nucleotide changes, yet they are largely -missed by the short-read data used in most large studies. +The structural variants (SVs) in this dataset were identified using Oxford +Nanopore long-read whole-genome sequencing of 888 individuals from the 1000 +Genomes Project, representing five ancestry groups. The SVs were merged with +previously identified short variants from the same individuals to generate a +multi-ancestry SV imputation reference panel. This panel was used to impute +SVs in approximately 500,000 UK Biobank participants and test their +associations with 32 disease-relevant traits.

-The panel contains more than 107,000 SVs called against GRCh38: about 60,000 -insertions, 38,000 deletions, 5,700 inversions, 2,700 breakends and 600 -duplications. Each variant carries an overall allele frequency and allele -frequencies for each of the five superpopulations (African, Admixed American, -East Asian, European, South Asian), Sniffles2 quality metrics, Hardy-Weinberg -p-values, and internal (leave-one-out) and UK Biobank imputation accuracy. -The authors used this panel to impute SVs into about 500,000 UK Biobank -participants and to test them for association with disease-relevant traits and -protein levels; where a variant reached genome-wide significance, the -associated traits are listed on its details page. +The track contains all 107,445 SVs in the reference panel: 59,953 insertions, +38,459 deletions, 5,729 inversions, 2,696 breakends, and 608 duplications. +Each variant is annotated with its overall allele frequency; allele +frequencies across five superpopulations (African, Admixed American, East +Asian, European, and South Asian); Hardy-Weinberg equilibrium p-values; and +imputation accuracy metrics from internal leave-one-out validation and UK +Biobank imputation. For SVs reaching genome-wide significance, the associated +traits, p-values, and INFO scores are listed on the corresponding variant +details page. +

+

+The 888 samples in this dataset are a subset of the 1,019 samples included in +the 1KG ONT Vienna track (Schloissnig et +al. 2025), with both datasets based on the same underlying sequencing data. +However, the data-processing and SV-calling methods differ between the two +tracks. The imputation reference panel, UK Biobank imputation results, and +SV-wide association study (SV-WAS) results described here are specific to this +track.

Display Conventions and Configuration

Items are colored by SV type, matching the other subtracks of the container:

@@ -41,31 +48,32 @@

Insertions and breakends are drawn at a single reference base; the length of inserted sequence is reported for insertions, and the mate locus of the rearrangement junction is reported for breakends. Deletions, inversions and duplications span the affected reference interval. Because the source table does not report an allele count, the allele count and allele number shown here are approximate values derived from the reported allele frequency and the genotype missing rate (allele number = 2 × 888 × (1 − missing rate); allele count = allele frequency × allele number).

The mouseover shows the variant name, SV type, reference and insertion lengths, allele frequency, approximate allele count, and the number of UK Biobank trait associations. Filters are available for SV type, SV length, insertion length, approximate allele count, overall and per-population allele frequency, the -number of UK Biobank GWAS hits, and the UK Biobank imputation r². +number of UK Biobank GWAS hits, and the leave-one-out imputation r² and +minor-allele concordance.

Methods

888 individuals from the 1000 Genomes Project (164 European, 144 Admixed American, 168 East Asian, 171 South Asian and 241 African), out of 906 sequenced, passed quality control. They were sequenced on the Oxford Nanopore PromethION P48 platform with R9.4.1 flow cells and the SQK-LSK110 ligation kit, to a median read length of about 6.2 kb and 15x median coverage. Reads were aligned to GRCh38 with minimap2 v2.24 and structural variants were jointly called across all samples with Sniffles2 v2.0.7 using tandem-repeat annotations. Variants were retained if they were 50 bp to 30 Mb long, present in at least two individuals and had a genotype missing rate below 20%, yielding 107,445 SVs. This SV panel was merged with about 45 million short variants from 1000 Genomes Phase 3 and phased with Beagle to build a multi-ancestry

  Deletion (DEL)
  Insertion (INS)
  Duplication (DUP)
  Inversion (INV)