29c46a47cbb40f44c06103e8294e8864d82256a7 max Mon Aug 17 05:56:59 2026 -0700 lrSv: fix typos and HTML consistency in track description pages GitHub capitalization, Continuous, missing article/period in HPRC2 row, 1000 Genomes capitalization, quote target/id attributes diff --git src/hg/makeDb/trackDb/human/lrSv1kLin.html src/hg/makeDb/trackDb/human/lrSv1kLin.html index 21fddf920bd..43c86e74399 100644 --- src/hg/makeDb/trackDb/human/lrSv1kLin.html +++ src/hg/makeDb/trackDb/human/lrSv1kLin.html @@ -1,154 +1,154 @@

Description

This track shows structural variants (SVs) from an integrated long-read callset spanning 1,218 individuals of the 1000 Genomes Project. Structural variants are genomic rearrangements larger than about 50 bp, such as deletions and insertions; because they alter large stretches of DNA at once, they can affect gene dosage and regulation more strongly than single-nucleotide changes, and long reads resolve them far better than short-read data.

Rather than coming from a single sequencing run, the calls are drawn together from several 1000 Genomes long-read efforts that use different technologies. The 1,218 individuals combine:

The current release contains more than 580,000 SVs on GRCh38 (391,410 insertions and 196,369 deletions) and more than 610,000 SVs on T2T-CHM13 (376,117 insertions and 238,405 deletions), each annotated with an overall allele frequency and allele frequencies for the five 1000 Genomes superpopulations (African, Admixed American, East Asian, European, South Asian).

Display Conventions and Configuration

Items are colored by SV type, matching the other subtracks of the container:

  Deletion (DEL)
  Insertion (INS)

Insertions are drawn at the insertion site with a width of 1 bp, and the length of inserted sequence is reported as the insertion length; deletions span the affected reference interval. The mouseover shows the variant name, SV type, reference and insertion lengths, allele count and per-population allele frequencies. Filters are available for SV type, SV length, insertion length, allele count, and overall and per-population allele frequency.

Methods

Lin et al. built the callset in two tiers. A baseline set of euchromatic SVs was established from 293 near-T2T haplotype-resolved assemblies generated by HGSVC and HPRC, and this was expanded with an additional 445 low-pass Oxford Nanopore genomes (Schloissnig et al. 2025) and 480 roughly 30x Oxford Nanopore genomes sequenced by the 1000 Genomes Long Read Sequencing Consortium (383 new plus 97 from Gustafson et al. 2024), for 1,218 genomes of diverse ancestry. Structural variants were discovered with ten long-read callers, and a machine-learning tool, BoostSV, ranked and selected the best allele to represent each SV across the different platforms and coverages, producing a single nonredundant callset. The final callset contains 614,522 SVs (376,117 insertions and 238,405 deletions) on T2T-CHM13 and 587,779 SVs (391,410 insertions and 196,369 deletions) on GRCh38. See Lin et al. for full details.

The insertion/deletion callset VCFs (GRCh38 and T2T-CHM13 native), already annotated with overall and per-superpopulation allele frequencies (EUR, AMR, EAS, AFR, SAS), were provided by the laboratories of Evan Eichler and Danny Miller (University of Washington). At UCSC the deletion and insertion records were converted to bigBed; no re-merging or re-annotation was performed. The step-by-step build commands (format conversion and bigBed build) are recorded in the UCSC makeDoc for this track container: doc/hg38/lrSv.txt. The conversion script and autoSql schema live in makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.

Data Access

The source Lin et al. dataset is available -from Github.

+from GitHub.

The data on the Genome Browser can be explored interactively in table format with the Table Browser or the Data Integrator and exported from there to spreadsheet or tab-separated tables. From scripts, the data can be accessed through our API, track=lrSv1kLin.

For automated download and analysis, the annotation is stored in bigBed files that can be downloaded from our download server: http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/lin1218.bb (GRCh38/hg38, native) and http://hgdownload.soe.ucsc.edu/gbdb/hs1/lrSv/lin1218.bb (T2T-CHM13/hs1, native). Individual regions or the whole annotation can be obtained with the bigBedToBed utility, which can be compiled from source or downloaded as a precompiled binary from our utilities page. The tool can also extract features within a given range, for example: bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/lin1218.bb -chrom=chr21 -start=0 -end=100000000 stdout.

Credits

Thanks to Jiadong Lin for providing the merged callset and the dataset overview, and to Evan Eichler, Danny Miller and colleagues at the University of Washington, together with the contributing 1000 Genomes long-read consortia (HPRC, HGSVC and the 1000 Genomes ONT sequencing groups), for generating and sharing this callset.

References

Lin J, et al. A high-resolution human pangenome structural variant resource for improved disease association. Submitted.

Schloissnig S, Pani S, Ebler J, Hain C, Tsapalou V, Söylev A, Hüther P, Ashraf H, Prodanov T, Asparuhova M et al. Structural variation in 1,019 diverse humans based on long-read sequencing. Nature. 2025 Aug;644(8076):442-452. PMID: 40702182; PMC: PMC12350158

Gustafson JA, Gibson SB, Damaraju N, Zalusky MPG, Hoekzema K, Twesigomwe D, Yang L, Snead AA, Richmond PA, De Coster W et al. High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive catalog of human genetic variation. Genome Res. 2024 Nov 20;34(11):2061-2073. PMID: 39358015; PMC: PMC11610458

Logsdon GA, Ebert P, Audano PA, Loftus M, Porubsky D, Ebler J, Yilmaz F, Hallast P, Prodanov T, Yoo D et al. Complex genetic variation in nearly complete human genomes. Nature. 2025 Aug;644(8076):430-441. PMID: 40702183; PMC: PMC12350169