95208355e2c667d194b29ee78c8ca8a09c2c2596 max Fri Jul 17 08:51:28 2026 -0700 lrSv: add NIH CARD long-read SV subtrack (cardSv) #Preview2 week - bugs introduced now will need a build patch to fix Add the NIH CARD Long-Read Initiative structural-variant catalogue (351 post-mortem brain samples: 205 NABEC European ancestry, 146 HBCC African/African-admixed) as a new subtrack of the Long-read SVs container. The provider bigBed is re-derived into the shared lrSv schema: signed svLen made positive (reference span), an explicit insLen added, the single DUP:TANDEM folded to DUP, and colors remapped to the container's shared svColor() palette. All 228,855 provider records are carried through 1:1. Adds the converter and autoSql, the trackDb stanza with filters consistent with the sibling subtracks, a full description page, a summary row and blurb on the container page, and a makeDoc section. refs #36258 diff --git src/hg/makeDb/trackDb/human/lrSv.html src/hg/makeDb/trackDb/human/lrSv.html index 73ecedd397d..615e2c4fd8e 100644 --- src/hg/makeDb/trackDb/human/lrSv.html +++ src/hg/makeDb/trackDb/human/lrSv.html @@ -192,30 +192,41 @@
Note: there is likely some overlap in sample composition across these collections. For example, 1000 Genomes samples are also included in HPRC and CoLoRSdb.
Structural variants from the Consortium of Long-Read Sequencing database (CoLoRSdb), from 1,427 PacBio HiFi long-read whole-genome sequences. ~426k SVs (insertions, deletions, inversions) called with pbsv and merged with Jasmine, with allele frequencies, genotype counts and Hardy-Weinberg statistics across the cohort.
@@ -334,30 +345,42 @@Structural variants from 101 long-read whole-genome sequences released alongside the GWAS SVatalog tool (Chirmade et al. 2026). The samples come from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT), a cystic-fibrosis (CF) patient cohort assembled to model patient-specific responses to CFTR modulator therapies (most participants are F508del homozygotes or F508del / minimal-function compound heterozygotes; a smaller number carry rare nonsense or missense CFTR mutations). ~87k SVs (deletions, insertions, duplications, inversions and complex events) annotated with gene overlaps, ClinGen / gnomAD constraint scores, OMIM / ClinVar / DGV / Decipher regional annotations.
++Structural variants from Oxford Nanopore long-read sequencing of post-mortem +brain tissue (prefrontal cortex) from 351 neurologically normal individuals, +generated by the NIH Center for Alzheimer's and Related Dementias (NIH CARD) +Long-Read Initiative (Billingsley et al. 2024). The cohort combines 205 +European-ancestry samples (North American Brain Expression Consortium, NABEC) +and 146 African / African-admixed samples (NIMH Human Brain Collection Core, +HBCC). ~229k SVs (insertions, deletions, inversions) with per-cohort carrier +counts and allele frequencies. +
+Each subtrack has its own documentation page with details on how to download and intersect the underlying annotations. The build process for all subtracks is recorded in the UCSC makeDoc, doc/hg38/lrSv.txt (and doc/hs1/lrSv.txt for T2T-CHM13); the conversion scripts are in makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.
Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, Corvelo A, Clarke WE, Musunuri R, Nagulapalli K et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 2022 Sep 1;185(18):3426-3440.e19. PMID: 36055201; PMC: PMC9439720
+ + ++Billingsley KJ, Meredith M, Daida K, Jerez PA, Negi S, Malik L, Genner RM, Moller A, Zheng X, Gibson +SB et al. + +Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural +variation on gene expression and DNA methylation. +bioRxiv. 2024 Dec 17;. +PMID: 39764002; PMC: PMC11702628 +
+