95208355e2c667d194b29ee78c8ca8a09c2c2596 max Fri Jul 17 08:51:28 2026 -0700 lrSv: add NIH CARD long-read SV subtrack (cardSv) #Preview2 week - bugs introduced now will need a build patch to fix Add the NIH CARD Long-Read Initiative structural-variant catalogue (351 post-mortem brain samples: 205 NABEC European ancestry, 146 HBCC African/African-admixed) as a new subtrack of the Long-read SVs container. The provider bigBed is re-derived into the shared lrSv schema: signed svLen made positive (reference span), an explicit insLen added, the single DUP:TANDEM folded to DUP, and colors remapped to the container's shared svColor() palette. All 228,855 provider records are carried through 1:1. Adds the converter and autoSql, the trackDb stanza with filters consistent with the sibling subtracks, a full description page, a summary row and blurb on the container page, and a makeDoc section. refs #36258 diff --git src/hg/makeDb/trackDb/human/cardSv.html src/hg/makeDb/trackDb/human/cardSv.html new file mode 100644 index 00000000000..233ee84f1b0 --- /dev/null +++ src/hg/makeDb/trackDb/human/cardSv.html @@ -0,0 +1,139 @@ +<h2>Description</h2> +<p> +This track shows structural variants (SVs) identified by Oxford Nanopore +long-read sequencing of post-mortem brain tissue (prefrontal cortex) from +351 individuals, generated by the NIH Center for Alzheimer's and Related +Dementias (NIH CARD) Long-Read Initiative. Structural variants are genomic +rearrangements larger than about 50 bp, such as deletions, insertions, +inversions and duplications; because they alter or move large stretches of +DNA at once they can have outsized effects on gene dosage, gene regulation +and DNA methylation compared with single-nucleotide changes. +</p> +<p> +The cohort combines two studies: 205 samples of European ancestry from the +North American Brain Expression Consortium (NABEC, dbGaP phs001300) and 146 +samples of African and African-admixed ancestry from the NIMH Human Brain +Collection Core (HBCC, dbGaP phs000979). The track contains more than 228,000 SVs called +against GRCh38: about 127,000 insertions, 102,000 deletions, 431 inversions +and one tandem duplication. Each record carries the number of carrier samples +overall and split by cohort (NABEC and HBCC), together with the allele frequency reported +by the source project. +</p> + +<h2>Display Conventions and Configuration</h2> +<p> +Items are colored by SV type, matching the other subtracks of the container: +<ul> +<li><span style="color: rgb(200,0,0);">Deletions (DEL)</span> - red</li> +<li><span style="color: rgb(0,0,200);">Insertions (INS)</span> - blue</li> +<li><span style="color: rgb(0,160,0);">Duplications (DUP)</span> - green</li> +<li><span style="color: rgb(230,140,0);">Inversions (INV)</span> - orange</li> +</ul> +</p> +<p> +Insertions are placed at the insertion site with a width of 1 bp, and the +length of the inserted sequence is shown as the insertion length; deletions, +inversions and the duplication span the affected reference interval. The +mouseover shows the variant name, SV type, reference and insertion lengths, +allele frequency and the number of carrier samples split into the NABEC and +HBCC cohorts. Filters are available for SV type, SV length, insertion length, +carrier count, allele frequency, and the number of carriers in each cohort. +</p> + +<h2>Methods</h2> +<p> +NABEC samples were sequenced on Oxford Nanopore R9.4.1 and HBCC samples on +R10.4.1 PromethION flow cells, with a median read N50 of 27 kb and about 40x +average genome coverage. Structural variants were called both from read +alignments (minimap2 alignments processed with Sniffles2 v2.3) and from +de novo assemblies (Shasta v0.11.1 assemblies phased with HapDup v0.12 and +compared to the reference with Hapdiff). Assembly-based calls were merged +across samples with Truvari, read-based calls were merged across samples with +Sniffles2, and the read and assembly sets were then merged together and across +the two cohorts with Truvari. All processing used the Nanopore Analysis +Pipeline (NAPU) workflows on the AnVIL/Terra platform; see Kolmogorov et al. +2023 and Billingsley et al. 2024 for details. +</p> +<p> +The display bigBed <tt>NIH_CARD_longReadSVs.bb</tt> was obtained from the NIH +CARD browser-track +<a href="https://github.com/meredith705/card_genome_browserTrack" target="_blank">GitHub +repository</a>. At UCSC it was converted to the shared long-read SV schema +(signed lengths made positive, an explicit insertion-length field added, the +single <tt>DUP:TANDEM</tt> call folded to <tt>DUP</tt>, and colors reassigned to +the container's shared palette) so it matches the other subtracks. The +step-by-step commands are recorded in the UCSC makeDoc for this track +container: +<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/lrSv.txt" target="_blank"> +doc/hg38/lrSv.txt</a>. The conversion script and autoSql schema live in +<a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/lrSv" target="_blank"> +makeDb/scripts/lrSv</a>, and the track configuration is in +<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/lrSv.ra" target="_blank">trackDb/human/lrSv.ra</a>. +</p> + +<h2>Data Access</h2> +<p> +The data can be explored interactively in table format with the +<a href="../cgi-bin/hgTables">Table Browser</a> or the +<a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there +to spreadsheet or tab-sep tables. From scripts, the data can be accessed +through our <a href="https://api.genome.ucsc.edu">API</a>, track=<i>cardSv</i>. +</p> +<p> +The annotation is stored as a bigBed file that can be downloaded from +<a href="http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/" target="_blank">our +download server</a> as <tt>card.bb</tt>. Individual regions or the whole +annotation can be obtained with the <tt>bigBedToBed</tt> utility, available +from our +<a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">utilities +page</a>. Example: +<tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/card.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt>. +</p> +<p> +The underlying genotype-level calls in VCF format are available under +controlled access through dbGaP (NABEC phs001300, substudy phs003181; HBCC +phs000979) and can be requested through the AnVIL Data Explorer for the +<a href="https://explore.anvilproject.org/datasets/0b740f36-28d6-4e02-8165-ad0e9674bbf6" target="_blank">NABEC</a> +and +<a href="https://explore.anvilproject.org/datasets/a4e936d1-d81a-475d-95be-b5cd41de921d" target="_blank">HBCC</a> +datasets. +</p> + +<h2>Credits</h2> +<p> +Thanks to the North American Brain Expression Consortium (NABEC), the NIMH +Human Brain Collection Core (HBCC), the Banner Sun Health Research Institute +Brain and Body Donation Program, and the NIH CARD Long-Read Initiative for +generating and sharing this dataset, and to Melissa Meredith for preparing the +browser track. This work was supported by the Intramural Research Programs of +the NIA, NINDS, NCI, NHGRI and NIMH, and used the NIH STRIDES Initiative and +the NIH HPC Biowulf cluster. +</p> + +<h2>References</h2> + + +<p> +Billingsley KJ, Meredith M, Daida K, Jerez PA, Negi S, Malik L, Genner RM, Moller A, Zheng X, Gibson +SB <em>et al</em>. +<a href="https://doi.org/10.1101/2024.12.16.628723" target="_blank"> +Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural +variation on gene expression and DNA methylation</a>. +<em>bioRxiv</em>. 2024 Dec 17;. +PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/39764002" target="_blank">39764002</a>; PMC: <a +href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11702628/" target="_blank">PMC11702628</a> +</p> + + + +<p> +Kolmogorov M, Billingsley KJ, Mastoras M, Meredith M, Monlong J, Lorig-Roach R, Asri M, Alvarez +Jerez P, Malik L, Dewan R <em>et al</em>. +<a href="https://doi.org/10.1038/s41592-023-01993-x" target="_blank"> +Scalable Nanopore sequencing of human genomes provides a comprehensive view of haplotype-resolved +variation and methylation</a>. +<em>Nat Methods</em>. 2023 Oct;20(10):1483-1492. +PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/37710018" target="_blank">37710018</a>; PMC: <a +href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11222905/" target="_blank">PMC11222905</a> +</p> +