23f68df8fb0481fd1bd4351f4adbf913f039f237 max Mon Jul 20 11:49:57 2026 -0700 lrSv1kLin: add per-population AF filter defaults and full Data Access section Add missing filter.afAfr/afAmr/afEas/afEur/afSas default ranges so the per-population allele-frequency filters render. Expand the description page's Data Access section to the standard form with bigBed download URLs for both the hg38 and hs1 native builds. refs #36258 diff --git src/hg/makeDb/trackDb/human/lrSv1kLin.html src/hg/makeDb/trackDb/human/lrSv1kLin.html index 9b819296a36..05b20758df6 100644 --- src/hg/makeDb/trackDb/human/lrSv1kLin.html +++ src/hg/makeDb/trackDb/human/lrSv1kLin.html @@ -1,95 +1,114 @@
This track shows structural variants (SVs) from an integrated long-read callset spanning 1,218 individuals of the 1000 Genomes Project. Structural variants are genomic rearrangements larger than about 50 bp, such as deletions and insertions; because they alter large stretches of DNA at once, they can affect gene dosage and regulation more strongly than single-nucleotide changes, and long reads resolve them far better than short-read data.
Rather than coming from a single sequencing run, the calls are drawn together from several 1000 Genomes long-read efforts that use different technologies: HiFi and genome-assembly-based calls from the Human Pangenome Reference Consortium (HPRC year 2), assembly-based calls from the Human Genome Structural Variation Consortium (HGSVC3), and Oxford Nanopore calls from the Vienna 1000 Genomes ONT release, together with Oxford Nanopore sequencing from the University of Washington 1000 Genomes ONT effort (see 1KG ONT UW). Sequencing of the 1000 Genomes collection is ongoing, so the number of individuals and variants in this track is expected to grow over time.
This track is preliminary and unpublished; its sample composition and variant counts will be updated as more long-read data becomes available.
The current release contains more than 580,000 SVs on GRCh38 (about 391,000 insertions and 196,000 deletions), each annotated with an overall allele frequency and allele frequencies for the five 1000 Genomes superpopulations (African, Admixed American, East Asian, European, South Asian). This is a preliminary, unpublished callset; the counts and sample composition will be updated as more data is added.
Items are colored by SV type, matching the other subtracks of the container:
| Deletion (DEL) | |
| Insertion (INS) |
Insertions are drawn at the insertion site with a width of 1 bp, and the length of inserted sequence is reported as the insertion length; deletions span the affected reference interval. The mouseover shows the variant name, SV type, reference and insertion lengths, allele count and per-population allele frequencies. Filters are available for SV type, SV length, insertion length, allele count, and overall and per-population allele frequency.
Per-sample long-read SV calls from the contributing 1000 Genomes efforts (HiFi and assembly-based calls from HPRC year 2 and HGSVC3, and Oxford Nanopore calls from the Vienna and University of Washington releases) were combined across the 1,218 individuals and merged into a single site-level callset with Truvari v5.2.0. Overall and per-superpopulation allele frequencies (EUR, AMR, EAS, AFR, SAS) were then added with bcftools fill-tags. Only deletions and insertions are reported in the current release. The callset is provided on both GRCh38/hg38 and T2T-CHM13/hs1 from the respective native assemblies.
The data was provided by the laboratories of Evan Eichler and Danny Miller (University of Washington) and is preliminary and unpublished; a manuscript is in preparation. The step-by-step build commands (format conversion and bigBed build) are recorded in the UCSC makeDoc for this track container: doc/hg38/lrSv.txt. The conversion script and autoSql schema live in makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.
The data can be explored interactively in table format with the Table Browser or the -Data Integrator, and accessed -programmatically through our API, -track=lrSv1kLin. +Data Integrator and exported from there +to spreadsheet or tab-separated tables. From scripts, the data can be accessed +through our API, track=lrSv1kLin. +
++For automated download and analysis, the annotation is stored in bigBed files +that can be downloaded from our download server: + +http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/lin1218.bb (GRCh38/hg38, +native) and + +http://hgdownload.soe.ucsc.edu/gbdb/hs1/lrSv/lin1218.bb (T2T-CHM13/hs1, +native). Individual regions or the whole annotation can be obtained with the +bigBedToBed utility, which can be compiled from source or downloaded +as a precompiled binary from our +utilities +page. The tool can also extract features within a given range, for example: +bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/lin1218.bb -chrom=chr21 -start=0 -end=100000000 stdout. +
++This is a preliminary, unpublished callset provided by the authors; the +underlying per-sample sequencing data is not yet publicly released.
Thanks to Evan Eichler, Danny Miller and colleagues at the University of Washington, and to the contributing 1000 Genomes long-read consortia (HPRC, HGSVC and the 1000 Genomes ONT sequencing groups), for generating and sharing this callset.