44c00f07b0e94306e09f30c84ea6ab0f044e1a29 max Fri Aug 14 05:09:12 2026 -0700 adding lin et al long-read SV subtrack, refs #38099 diff --git src/hg/makeDb/trackDb/human/lrSv.html src/hg/makeDb/trackDb/human/lrSv.html index c2ad0693e4a..4901b321bcc 100644 --- src/hg/makeDb/trackDb/human/lrSv.html +++ src/hg/makeDb/trackDb/human/lrSv.html @@ -1,565 +1,665 @@

Description

This track collection contains structural variant (SV) calls derived from long-read sequencing studies. Structural variants are genomic rearrangements larger than ~50 bp, including deletions, insertions, duplications, inversions, and translocations. Long-read sequencing technologies can span repetitive regions and resolve complex rearrangements that are difficult to detect with short-read methods. -

+The long read datasets described below were produced with one of two sequencing technologies, Oxford Nanopore +Technologies (ONT) or Pacific Biosciences (PacBio, whose highly accurate reads are also called +HiFi, unlike the longer but less accurate CLR, Continous Long-Reads).

Available Datasets

-SV length statistics (min / median / max) are computed from the svLen -field of each track, in base pairs. Some tracks include sites with -svLen=0 (complex events where the reference and alternate alleles -differ in sequence but not in length). +SV length statistics (min / median / max) use the size of the variant in base +pairs: the inserted-sequence length for insertions and the reference span for +deletions and other types. (For insertions the svLen reference-span +field is only a 1-2 bp placeholder, so the inserted length is reported +instead.) Some tracks include sites of length 0, complex events where the +reference and alternate alleles differ in sequence but not in length. +For example, two different insertions, one in either sequence, is usually +called a "complex" event.

-

+

Polymorphic Mobile Element Insertions (Alu, L1, SVA, HERVK, snRNA) called from HGSVC3 long-read assemblies are released as a separate track collection; see the Mobile Insertions tracks. Those MEIs are the insertions identified in the 65 HGSVC3 samples relative to the reference, available on both GRCh38/hg38 and T2T-CHM13/hs1.

- - + + - - - - - - - - - + + + + + + + + + - - - + + + - - - - - + + + + + - + - - - + + + - - - - - - - - - - - - - - - - - - - - - - - - - - - + + + + + - - - + + + - - + + - - + + - + + + + + + + + + + + + - - - + + + - - - + + + + + + + + + + + + + + - + - - - + + + - - + + - - + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + - - - - - - - - - - - - - - - - - - - - - -
Dataset N samples Cohort / disease Disease cases Coverage SV count Min Median Max
All merged All long-read SV datasets merged on identical position+type+length, with per-database AC mixedmixed (PacBio HiFi, ONT)2,582,278mixed (HiFi, ONT)3,111,026 1 1 57,207,413
CoLoRSdb 1,427 Consortium of Long-Read Sequencing, joint callset No mixed (HiFi) 426,239 20 33 101,381
Han 945945Han Chinese, general populationNo~17x ONT111,288125499,744AoU 1K1,027All of Us, self-identified Black/African American; biobank includes a variety of conditions (diabetes, hearing loss, etc.), filtered for allele count > 20Yes (mixed)~8x HiFi540,155501529,998
1KG ONT UW1001000 Genomes, 5 superpopulations / 19 subpopulations (University of Washington ONT effort)1KG Lin merged1,2181000 Genomes long-read merge (HiFi assemblies + ONT), Lin et al. No~37x ONT (R9.4.1)113,159116798,290mixed (HiFi assembly, ONT R9/R10)587,7795017199,968
1KG ONT Vienna1KG Vienna ONT 1,019 1000 Genomes, diverse No ~17x ONT 148,375 2 157 49,171
ToMMo Japanese333 (111 trios)Japanese, general population1KG UW ONT1001000 Genomes, 5 superpopulations / 19 subpopulations (University of Washington ONT effort) No~22x ONT74,2015115899,985
AoU 1K1,027All of Us, self-identified Black/African American; biobank includes a variety of conditions (diabetes, hearing loss, etc.)Yes (mixed)~8x HiFi540,155501529,998
GA4K502Children's Mercy, pediatric rare disease probands + familiesYes (probands)~27x HiFi115,55450186809,712~37x ONT (R9.4.1)113,159116798,290
deCODE 3,6223,622Icelandic general population1KG Boehringer ONT 8888881000 Genomes, 5 superpopulations; used to impute SVs into ~500,000 UK Biobank participants No~17x ONT119,453~15x ONT (R9.4.1)107,445 1154861,081128,634,664
HPRC v2.1 233HPRC release-2 pangenome (CHM13 + diverse 1KG assemblies)HPRC 2.1 pangenome (CHM13, some 1KG assemblies) No ~60x HiFi + ~30x ONT (pangenome graph) 549,649 50 261 1,064,897
HGSVC365HGSVC3 diverse reference assembliesNo~47x HiFi + ~56x ONT176,5315015430,176,500
HGSVC2 32 HGSVC2 haplotype-resolved assemblies (5 superpopulations) No >40x PacBio CLR + >20x HiFi (+ Strand-seq) 111,746 50 168 57,207,413
HGSVC365HGSVC3 diverse reference assembliesNIH CARD 351351NIH CARD post-mortem brain (prefrontal cortex); NABEC (European) + HBCC (African/African-admixed), no Alzheimer's disease cases No~47x HiFi + ~56x ONT176,53150~40x ONT (R9.4.1 / R10.4.1)228,8551130,282,742
deCODE 3,6223,622Icelandic general populationNo~17x ONT119,4531 15430,176,500861,081
Arab APR53UAE-resident Arabs from 8 countries (Arab Pangenome Reference)Han 945945Han Chinese, general population No~35x HiFi + ~54x ONT (+ Hi-C, pangenome graph)72,656~17x ONT111,288 1121584,01625499,744
CPC 58 Chinese Pangenome Consortium, 36 minority ethnic groups (HPRC-specific SVs removed) No ~30x HiFi (pangenome graph) 36,030 50 134 8,998,096
ToMMo Japanese333 (111 trios)Japanese, general populationNo~22x ONT74,2015115899,985
Arab APR53UAE-resident Arabs from 8 countries (Arab Pangenome Reference)No~35x HiFi + ~54x ONT (+ Hi-C, pangenome graph)72,6561121584,016
GA4K502Children's Mercy, pediatric rare disease probands + familiesYes (probands)~27x HiFi115,55450186809,712
SVatalog 101 101 Cystic fibrosis (CF) patients from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT). Long-read WGS used for GWAS LD fine-mapping Yes (all CF) ~50x PacBio CLR (34, Sequel I) + ~76x HiFi (67, Sequel II) 87,068 4 160 1,321,484
NIH CARD 351351NIH CARD post-mortem brain (prefrontal cortex); NABEC (European) + HBCC (African/African-admixed), no Alzheimer's disease casesNo~40x ONT (R9.4.1 / R10.4.1)228,8551130,282,742
Noyvert 8888881000 Genomes, 5 superpopulations; used to impute SVs into ~500,000 UK Biobank participantsNo~15x ONT (R9.4.1)107,4451128,634,664

Note: there is likely some overlap in sample composition across these collections. For example, 1000 Genomes samples are also included in HPRC and CoLoRSdb.

+

1000 Genomes long-read callsets

+

+Several of the datasets above are long-read callsets on the 1000 Genomes +Project samples, produced by different groups with different technologies and +variant-calling strategies. The +1KG Lin merged track combines these and added more 1000 genomes assemblies for a +single 1,218-individual callset (Lin et al., submitted). The table below lists +the merged release and the callsets that contribute to it. +

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
CallsetN samplesStudyData sourceVariant callingUCSC track
Lin_12181,218Lin et al., A high-resolution human pangenome structural variant resource for improved disease associationHGSVC3, HPRC2 293 graph+linear, UW ONT (480; 383 newly generated by UW, 97 from Gustafson et al.), Vienna ONT (445)Linear-reference-based merge1KG Lin merged
Vienna ONT1,019Schloissnig et al., Structural variation in 1,019 diverse humans based on long-read sequencingLow-pass 17x ONTGraph-based1KG Vienna ONT
UW ONT100Gustafson et al., High-coverage nanopore sequencing of samples from the 1000 Genomes ProjectHigh-coverage 37x ONTLinear-reference-based1KG UW ONT
HPRC2 Minigraph-Cactus232Lucas et al., HPRC2: a human pangenome reference with near-complete coverage of common genetic variationNear-T2T assembly (30x-60x HiFi+ONT), a linear callset was included for Lin et al merge. Minigraph-cactus graphHPRC v2.1
HGSVC365Logsdon et al., Complex genetic variation in nearly complete human genomesNear-T2T assembly (37-40x HiFi+ONT)Linear-reference-basedHGSVC3
+

CoLoRSdb SVs

Structural variants from the Consortium of Long-Read Sequencing database (CoLoRSdb), from 1,427 PacBio HiFi long-read whole-genome sequences. ~426k SVs (insertions, deletions, inversions) called with pbsv and merged with Jasmine, with allele frequencies, genotype counts and Hardy-Weinberg statistics across the cohort.

-

Han 945 SVs

+

AoU 1K SVs

-Structural variants from 945 Han Chinese individuals. ~111k SVs -(deletions, insertions, duplications, inversions, translocations) merged with SURVIVOR. -Includes allele frequencies and per-sample support. +Structural variants from 1,027 individuals from the All of Us (AoU) Research Program, +sequenced with PacBio HiFi long reads. AoU is a deeply phenotyped biobank +that includes participants with a range of conditions (e.g. diabetes, +hearing loss, hypertension), so the cohort is not disease-free. +~541k SVs (insertions and deletions) with population-specific allele +frequencies, gene annotations, and clinical trait associations. +Due to data sharing rules of the All of Us project, only variants with an allele +count > 20 can be shown here.

-

1KG ONT UW SVs

+

1KG Lin merged SVs

-Structural variants from Oxford Nanopore long-read sequencing of 100 -1000 Genomes samples (5 superpopulations, 19 subpopulations) from the -University of Washington-led 1000 Genomes ONT sequencing effort, described in -Gustafson et al. 2024. ~114k SVs (insertions, deletions, duplications, -inversions) called with five callers and merged with Jasmine. This is mostly a -separate dataset from the Vienna 1KG-ONT release described next (directly below); -only two samples (HG03499 and HG03548) overlap. +A single nonredundant long-read callset across 1,218 1000 Genomes individuals +(Lin et al.), combining 293 near-T2T HGSVC and HPRC assemblies, 480 University +of Washington Oxford Nanopore genomes (including the Gustafson et al. samples), +and 445 low-pass Oxford Nanopore genomes from the Vienna release. Structural +variants from ten long-read callers were integrated with the BoostSV +machine-learning tool. ~588k SVs on GRCh38 (391k insertions, 196k deletions), +each with an overall and per-superpopulation allele frequency. Because it +already merges several of the 1000 Genomes datasets above, its samples are also +represented in those individual tracks.

-

1KG ONT Vienna SVs

+

1KG Vienna ONT SVs

Structural variants from 1,019 individuals across 26 populations (1000 Genomes ONT). ~161k SVs annotated with SVAN, classifying insertions and deletions by mechanism of origin (mobile elements, VNTRs, processed pseudogenes, etc.). Original coordinates are on T2T-CHM13 (hs1); the hg38 version was created via liftOver. -Two samples (HG03499 and HG03548) overlap with the 1KG ONT UW dataset. +Two samples (HG03499 and HG03548) overlap with the 1KG UW ONT dataset.

-

ToMMo Japanese SVs

+

1KG UW ONT SVs

-Structural variants from 333 Japanese individuals (111 trios) from the Tohoku Medical -Megabank (ToMMo). ~74k SVs (deletions and insertions) with trio-based Mendelian -error rates and allele frequencies. -

- -

AoU 1K SVs

-

-Structural variants from 1,027 individuals from the All of Us (AoU) Research Program, -sequenced with PacBio HiFi long reads. AoU is a deeply phenotyped biobank -that includes participants with a range of conditions (e.g. diabetes, -hearing loss, hypertension), so the cohort is not disease-free. -~541k SVs (insertions and deletions) with population-specific allele -frequencies, gene annotations, and clinical trait associations. +Structural variants from Oxford Nanopore long-read sequencing of 100 +1000 Genomes samples (5 superpopulations, 19 subpopulations) from the +University of Washington-led 1000 Genomes ONT sequencing effort, described in +Gustafson et al. 2024. ~114k SVs (insertions, deletions, duplications, +inversions) called with five callers and merged with Jasmine. This is mostly a +separate dataset from the 1KG Vienna ONT release described directly above; +only two samples (HG03499 and HG03548) overlap.

-

GA4K SVs

+

1KG Boehringer ONT 888 SVs

-Structural variants from 502 probands and family members enrolled in the -Genomic Answers for Kids (GA4K) pediatric rare-disease program at Children's -Mercy Research Institute, sequenced with PacBio HiFi long reads. ~116k -replicated SVs (deletions, insertions, duplications, inversions) called with -pbsv and merged with JASMINE. The matched GA4K small-variant callset (SNVs -and short indels) lives alongside other population allele-frequency resources -as GA4K 552 PacBio LR in the Variant -Frequencies track collection. +Structural variants from Oxford Nanopore long-read sequencing of 888 +individuals from the 1000 Genomes Project, spanning five ancestry groups +(European, Admixed American, East Asian, South Asian, African; Noyvert et al. +2025). ~107k SVs (insertions, deletions, inversions, breakends and +duplications) called with Sniffles2, with overall and per-superpopulation +allele frequencies, Sniffles2 and Hardy-Weinberg quality metrics, and +imputation accuracy. The panel was used to impute SVs into about 500,000 UK +Biobank participants and test them for association with disease traits and +protein levels; genome-wide significant UK Biobank associations are listed on +each variant's details page.

-

deCODE 3,622 SVs

-

-High-confidence structural variants from 3,622 Icelanders (deCODE genetics), -sequenced with Oxford Nanopore long reads. ~134k SVs (deletions, insertions -and combined insertion/deletion events). Site-only callset with annotated -surrounding tandem-repeat regions. -

HPRC v2.1 SVs

Structural variants derived from the Human Pangenome Reference Consortium release-2.1 minigraph-cactus pangenome graph, built from 233 PacBio HiFi haplotype-resolved assemblies (CHM13 + diverse 1000 Genomes samples). About 550k SV-sized alleles (insertions and deletions) extracted from the -graph with vg deconstruct. +graph with vg deconstruct. A more traditional, linear callset was +made by Wenwei Liao and is available +from GitHub. +

+ +

HGSVC3 65 SVs

+

+Structural variants from 65 diverse individuals sequenced and de novo +assembled by the Human Genome Structural Variation Consortium phase 3 +(HGSVC3). ~177k haplotype-resolved SVs (deletions, insertions and +inversions) called with PAV and cross-validated with ten additional callers, +with per-site carrier haplotype lists and structural annotations.

HGSVC2 32 SVs

Structural variants from 32 haplotype-resolved diploid genomes (HGSVC2 freeze 4, Ebert et al. 2021). ~112k SVs (deletions, insertions and inversions) called from phased de novo assemblies with PAV, with per-variant 1000 Genomes population allele frequencies (insertions and deletions) and rich structural/gene annotations. An earlier HGSVC release complementary to HGSVC3.

-

HGSVC3 65 SVs

+

NIH CARD 351 SVs

-Structural variants from 65 diverse individuals sequenced and de novo -assembled by the Human Genome Structural Variation Consortium phase 3 -(HGSVC3). ~177k haplotype-resolved SVs (deletions, insertions and -inversions) called with PAV and cross-validated with ten additional callers, -with per-site carrier haplotype lists and structural annotations. +Structural variants from Oxford Nanopore long-read sequencing of post-mortem +brain tissue (prefrontal cortex) from 351 individuals, generated by the NIH +Center for Alzheimer's and Related Dementias (NIH CARD) Long-Read Initiative +(Billingsley et al. 2024). These are population brain-tissue cohorts with no +Alzheimer's disease cases. The cohort combines 205 European-ancestry samples +(North American Brain Expression Consortium, NABEC) and 146 African / +African-admixed samples (NIMH Human Brain Collection Core, HBCC). ~229k SVs +(insertions, deletions, inversions) with per-cohort allele counts and allele +frequencies. +

+ +

deCODE 3,622 SVs

+

+High-confidence structural variants from 3,622 Icelanders (deCODE genetics), +sequenced with Oxford Nanopore long reads. ~134k SVs (deletions, insertions +and combined insertion/deletion events). Site-only callset with annotated +surrounding tandem-repeat regions. +

+ +

Han 945 SVs

+

+Structural variants from 945 Han Chinese individuals. ~111k SVs +(deletions, insertions, duplications, inversions, translocations) merged with SURVIVOR. +Includes allele frequencies and per-sample support. +

+ +

CPC 58 SVs

+

+Structural variants from the Chinese Pangenome Consortium (CPC), 58 samples +spanning 36 minority ethnic groups (PacBio HiFi pangenome graph; Gao et al. +2023). This track shows the CPC contribution to the joint CPC+HPRC graph with +HPRC-specific SVs removed. ~36k SVs on hg38 (deletions, insertions and mixed +snarls), lifted from the native T2T-CHM13 assembly; the hs1 track is native. +

+ +

ToMMo Japanese SVs

+

+Structural variants from 333 Japanese individuals (111 trios) from the Tohoku Medical +Megabank (ToMMo). ~74k SVs (deletions and insertions) with trio-based Mendelian +error rates and allele frequencies.

Arab APR 53 SVs

Structural variants from the Arab Pangenome Reference (APR), a haplotype-resolved pangenome graph built from 53 UAE-resident Arab individuals drawn from eight countries (PacBio HiFi + ultralong ONT + Hi-C; Nassir et al. 2025). ~73k SVs on hg38 (deletions, insertions, complex and mixed snarls), lifted from the native T2T-CHM13 assembly; the hs1 track uses the native coordinates.

-

CPC 58 SVs

+

GA4K SVs

-Structural variants from the Chinese Pangenome Consortium (CPC), 58 samples -spanning 36 minority ethnic groups (PacBio HiFi pangenome graph; Gao et al. -2023). This track shows the CPC contribution to the joint CPC+HPRC graph with -HPRC-specific SVs removed. ~36k SVs on hg38 (deletions, insertions and mixed -snarls), lifted from the native T2T-CHM13 assembly; the hs1 track is native. +Structural variants from 502 probands and family members enrolled in the +Genomic Answers for Kids (GA4K) pediatric rare-disease program at Children's +Mercy Research Institute, sequenced with PacBio HiFi long reads. ~116k +replicated SVs (deletions, insertions, duplications, inversions) called with +pbsv and merged with JASMINE. The matched GA4K small-variant callset (SNVs +and short indels) lives alongside other population allele-frequency resources +as GA4K 552 PacBio LR in the Variant +Frequencies track collection.

SVatalog 101 SVs - Cystic Fibrosis

Structural variants from 101 long-read whole-genome sequences released alongside the GWAS SVatalog tool (Chirmade et al. 2026). The samples come from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT), a cystic-fibrosis (CF) patient cohort assembled to model patient-specific responses to CFTR modulator therapies (most participants are F508del homozygotes or F508del / minimal-function compound heterozygotes; a smaller number carry rare nonsense or missense CFTR mutations). ~87k SVs (deletions, insertions, duplications, inversions and complex events) annotated with gene overlaps, ClinGen / gnomAD constraint scores, OMIM / ClinVar / DGV / Decipher regional annotations.

-

NIH CARD 351 SVs

-

-Structural variants from Oxford Nanopore long-read sequencing of post-mortem -brain tissue (prefrontal cortex) from 351 individuals, generated by the NIH -Center for Alzheimer's and Related Dementias (NIH CARD) Long-Read Initiative -(Billingsley et al. 2024). These are population brain-tissue cohorts with no -Alzheimer's disease cases. The cohort combines 205 European-ancestry samples -(North American Brain Expression Consortium, NABEC) and 146 African / -African-admixed samples (NIMH Human Brain Collection Core, HBCC). ~229k SVs -(insertions, deletions, inversions) with per-cohort allele counts and allele -frequencies. -

- -

Noyvert 888 SVs

-

-Structural variants from Oxford Nanopore long-read sequencing of 888 -individuals from the 1000 Genomes Project, spanning five ancestry groups -(European, Admixed American, East Asian, South Asian, African; Noyvert et al. -2025). ~107k SVs (insertions, deletions, inversions, breakends and -duplications) called with Sniffles2, with overall and per-superpopulation -allele frequencies, Sniffles2 and Hardy-Weinberg quality metrics, and -imputation accuracy. The panel was used to impute SVs into about 500,000 UK -Biobank participants and test them for association with disease traits and -protein levels; genome-wide significant UK Biobank associations are listed on -each variant's details page. -

- -

Data Access

Each subtrack has its own documentation page with details on how to download and intersect the underlying annotations. The build process for all subtracks is recorded in the UCSC makeDoc, doc/hg38/lrSv.txt (and doc/hs1/lrSv.txt for T2T-CHM13); the conversion scripts are in makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.

References

+

+Lin J, et al. A high-resolution human pangenome structural variant +resource for improved disease association. Submitted. +(1KG Lin merged callset and dataset overview provided by Jiadong Lin.) +

+

Gong J, Sun H, Wang K, Zhao Y, Huang Y, Chen Q, Qiao H, Gao Y, Zhao J, Ling Y et al. Long-read sequencing of 945 Han individuals identifies structural variants associated with phenotypic diversity and disease susceptibility. Nat Commun. 2025 Feb 10;16(1):1494. PMID: 39929826; PMC: PMC11811171

Schloissnig S, Pani S, Ebler J, Hain C, Tsapalou V, Söylev A, Hüther P, Ashraf H, Prodanov T, Asparuhova M et al. Structural variation in 1,019 diverse humans based on long-read sequencing. Nature. 2025 Aug;644(8076):442-452. PMID: 40702182; PMC: PMC12350158

Otsuki A, Okamura Y, Ishida N, Tadaka S, Takayama J, Kumada K, Kawashima J, Taguchi K, Minegishi N, Kuriyama S et al. Construction of a trio-based structural variation panel utilizing activated T lymphocytes and long- read sequencing technology. Commun Biol. 2022 Sep 20;5(1):991. PMID: 36127505; PMC: PMC9489684

Garimella KV, Li Q, Wertz J, Lee SK, Cunial F, Huang Y, Mostovoy Y, Lorig-Roach R, English A, Su H et al. Population-scale Long-read Sequencing in the All of Us Research Program. medRxiv. 2025 Oct 5;. PMID: 41256123; PMC: PMC12622093

Cohen ASA, Farrow EG, Abdelmoity AT, Alaimo JT, Amudhavalli SM, Anderson JT, Bansal L, Bartik L, Baybayan P, Belden B et al. Genomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes. Genet Med. 2022 Jun;24(6):1336-1348. PMID: 35305867

Beyter D, Ingimundardottir H, Oddsson A, Eggertsson HP, Bjornsson E, Jonsson H, Atlason BA, Kristmundsdottir S, Mehringer S, Hardarson MT et al. Long-read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits. Nat Genet. 2021 Jun;53(6):779-786. PMID: 33972781

Logsdon GA, Ebert P, Audano PA, Loftus M, Porubsky D, Ebler J, Yilmaz F, Hallast P, Prodanov T, Yoo D et al. Complex genetic variation in nearly complete human genomes. Nature. 2025 Aug;644(8076):430-441. PMID: 40702183; PMC: PMC12350169

Chirmade S, Wang Z, Mastromatteo S, Sanders E, Thiruvahindrapuram B, Nalpathamkalam T, Pellecchia G, Lin F, Keenan K, Patel RV et al. GWAS SVatalog: a visualization tool to aid fine-mapping of GWAS loci with structural variations. Heredity (Edinb). 2026 Mar;135(3):199-210. PMID: 41203876; PMC: PMC13031531

Gustafson JA, Gibson SB, Damaraju N, Zalusky MPG, Hoekzema K, Twesigomwe D, Yang L, Snead AA, Richmond PA, De Coster W et al. High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive catalog of human genetic variation. Genome Res. 2024 Nov 20;34(11):2061-2073. PMID: 39358015; PMC: PMC11610458

Ebert P, Audano PA, Zhu Q, Rodriguez-Martin B, Porubsky D, Bonder MJ, Sulovari A, Ebler J, Zhou W, Serra Mari R et al. Haplotype-resolved diverse human genomes and integrated analysis of structural variation. Science. 2021 Apr 2;372(6537). PMID: 33632895; PMC: PMC8026704

Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, Corvelo A, Clarke WE, Musunuri R, Nagulapalli K et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 2022 Sep 1;185(18):3426-3440.e19. PMID: 36055201; PMC: PMC9439720

Billingsley KJ, Meredith M, Daida K, Jerez PA, Negi S, Malik L, Genner RM, Moller A, Zheng X, Gibson SB et al. Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural variation on gene expression and DNA methylation. bioRxiv. 2024 Dec 17;. PMID: 39764002; PMC: PMC11702628

Noyvert B, Erzurumluoglu AM, Drichel D, Omland S, Andlauer TFM et al. Imputation of structural variants using a multi-ancestry long-read sequencing panel enables identification of disease associations. eLife. 2025. doi:10.7554/eLife.106115.1