e11e10c01c975653b7f0102601cabd52967d2c80 max Fri Aug 14 05:58:23 2026 -0700 lrSv: author-provided noyvertSv description, Vienna ONT naming, hs1 Lin update, refs #38099 - noyvertSv.html: replace the Description with the author-provided text (imputation purpose, singletons excluded, subset-of-Vienna relationship) - rename "1KG ONT Vienna" -> "1KG Vienna ONT" to match the subtrack and merged-track labels (noyvertSv.html and the hs1 lrSv page) - hs1 lrSv page: add the 1KG Lin merged subtrack (now native on T2T-CHM13, 614,522 SVs) and reorder the summary table and detail sections to match the track (priority) order - lrSv1kLin.html: link the source Lin et al. dataset on GitHub diff --git src/hg/makeDb/trackDb/human/hs1/html/lrSv.html src/hg/makeDb/trackDb/human/hs1/html/lrSv.html index 3f3a28ece6e..5ddacee7dd7 100644 --- src/hg/makeDb/trackDb/human/hs1/html/lrSv.html +++ src/hg/makeDb/trackDb/human/hs1/html/lrSv.html @@ -1,116 +1,128 @@
This track collection shows structural variants (SVs) from long-read sequencing, called natively against the T2T-CHM13 reference (hs1). It is the T2T-CHM13 companion to the GRCh38 Long-read SVs collection.
Only a subset of the long-read SV datasets have been released with native T2T-CHM13 coordinates, and those are the subtracks shown here. The remaining cohorts in the collection were released on GRCh38 only. For the full set of datasets, see the hg38 Long-read SVs track.
| Dataset | N samples | Technology | SV count (hs1) |
|---|---|---|---|
| CoLoRSdb | 1,427 | PacBio HiFi | 839,714 |
| 1KG ONT Vienna | 1,019 | ONT | 161,332 |
| HGSVC3 | 65 | HiFi + ONT | 188,500 |
| 1KG Lin merged | 1,218 | ONT + assembly (merged) | 614,522 |
| 1KG Vienna ONT | 1,019 | ONT | 161,332 |
| HPRC v2.1 | 233 | Pangenome (minigraph-cactus) | 541,176 |
| Arab APR | 53 | HiFi + ONT pangenome | 103,077 |
| HGSVC3 | 65 | HiFi + ONT | 188,500 |
| CPC | 58 | HiFi pangenome | 46,092 |
| Arab APR | 53 | HiFi + ONT pangenome | 103,077 |
HGSVC3 and HPRC v2.1 are built directly from the consortia's T2T-CHM13 -releases; CoLoRSdb, 1KG ONT Vienna, Arab APR, and CPC are built from callsets -or pangenome graphs native to T2T-CHM13. Per-subtrack details, cohorts, and -citations are on each subtrack's own description page. +releases; CoLoRSdb, 1KG Lin merged, 1KG Vienna ONT, Arab APR, and CPC are built +from callsets or pangenome graphs native to T2T-CHM13. Per-subtrack details, +cohorts, and citations are on each subtrack's own description page.
Structural variants from the Consortium of Long-Read Sequencing database (CoLoRSdb), from 1,427 PacBio HiFi long-read whole-genome sequences. ~840k SVs (insertions, deletions, inversions) called with pbsv and merged with Jasmine, with allele frequencies, genotype counts and Hardy-Weinberg statistics across the cohort.
-+A merged long-read SV callset spanning 1,218 individuals of the 1000 Genomes +Project (Lin et al.), combining 293 near-T2T haplotype-resolved assemblies +(HPRC and HGSVC), 480 University of Washington Oxford Nanopore genomes, and 445 +Vienna Oxford Nanopore genomes. Structural variants were discovered with ten +long-read callers, and the BoostSV machine-learning tool selected the best +allele to represent each SV across platforms and coverages. ~615k SVs +(insertions and deletions) in native T2T-CHM13 coordinates. +
+ +Structural variants from 1,019 individuals across 26 populations (1000 Genomes ONT), called natively against T2T-CHM13. ~161k SVs annotated with SVAN, classifying insertions and deletions by mechanism of origin (mobile elements, VNTRs, processed pseudogenes, and others).
++Structural variants derived from the Human Pangenome Reference Consortium +release-2.1 minigraph-cactus pangenome graph, built from 233 PacBio HiFi +haplotype-resolved assemblies. ~541k SV-sized alleles (insertions and +deletions) extracted from the T2T-CHM13 graph with vg deconstruct. +
+Structural variants from 65 diverse individuals sequenced and de novo assembled by the Human Genome Structural Variation Consortium phase 3 (HGSVC3), from the consortium's native T2T-CHM13 annotation tables. ~189k haplotype-resolved SVs (deletions, insertions and inversions) called with PAV and cross-validated with ten additional callers, with per-site carrier haplotype lists and structural annotations.
--Structural variants derived from the Human Pangenome Reference Consortium -release-2.1 minigraph-cactus pangenome graph, built from 233 PacBio HiFi -haplotype-resolved assemblies. ~541k SV-sized alleles (insertions and -deletions) extracted from the T2T-CHM13 graph with vg deconstruct. +Structural variants from the Chinese Pangenome Consortium (CPC), 58 samples +spanning 36 minority ethnic groups (PacBio HiFi pangenome graph; Gao et al. +2023). This track shows the CPC contribution to the joint CPC+HPRC graph with +HPRC-specific SVs removed. ~46k SVs (deletions, insertions and mixed snarls) in +native T2T-CHM13 coordinates.
Structural variants from the Arab Pangenome Reference (APR), a haplotype-resolved pangenome graph built from 53 UAE-resident Arab individuals drawn from eight countries (PacBio HiFi + ultralong ONT + Hi-C; Nassir et al. 2025). ~103k SVs (deletions, insertions, complex and mixed snarls) in native T2T-CHM13 coordinates.
--Structural variants from the Chinese Pangenome Consortium (CPC), 58 samples -spanning 36 minority ethnic groups (PacBio HiFi pangenome graph; Gao et al. -2023). This track shows the CPC contribution to the joint CPC+HPRC graph with -HPRC-specific SVs removed. ~46k SVs (deletions, insertions and mixed snarls) in -native T2T-CHM13 coordinates. -
-Items are colored by SV type:
Each subtrack has its own documentation page with details on how to download and intersect the underlying annotations. The T2T-CHM13 build steps are recorded in the UCSC makeDoc, doc/hs1/lrSv.txt (with the shared pipeline in doc/hg38/lrSv.txt); the conversion scripts are in makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.
If you know of additional long-read structural-variant datasets on T2T-CHM13 that we could add, please contact us at genome@soe.ucsc.edu.