95208355e2c667d194b29ee78c8ca8a09c2c2596 max Fri Jul 17 08:51:28 2026 -0700 lrSv: add NIH CARD long-read SV subtrack (cardSv) #Preview2 week - bugs introduced now will need a build patch to fix Add the NIH CARD Long-Read Initiative structural-variant catalogue (351 post-mortem brain samples: 205 NABEC European ancestry, 146 HBCC African/African-admixed) as a new subtrack of the Long-read SVs container. The provider bigBed is re-derived into the shared lrSv schema: signed svLen made positive (reference span), an explicit insLen added, the single DUP:TANDEM folded to DUP, and colors remapped to the container's shared svColor() palette. All 228,855 provider records are carried through 1:1. Adds the converter and autoSql, the trackDb stanza with filters consistent with the sibling subtracks, a full description page, a summary row and blurb on the container page, and a makeDoc section. refs #36258 diff --git src/hg/makeDb/trackDb/human/lrSv.html src/hg/makeDb/trackDb/human/lrSv.html index 73ecedd397d..615e2c4fd8e 100644 --- src/hg/makeDb/trackDb/human/lrSv.html +++ src/hg/makeDb/trackDb/human/lrSv.html @@ -192,30 +192,41 @@ <td>50</td> <td>134</td> <td>8,998,096</td> </tr> <tr> <td><a href="hgTrackUi?g=chirmade101Sv">SVatalog 101</a></td> <td>101</td> <td>Cystic fibrosis (CF) patients from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT). Long-read WGS used for GWAS LD fine-mapping</td> <td>Yes (all CF)</td> <td>~50x PacBio CLR (34, Sequel I) + ~76x HiFi (67, Sequel II)</td> <td>87,068</td> <td>4</td> <td>160</td> <td>1,321,484</td> </tr> +<tr> + <td><a href="hgTrackUi?g=cardSv">NIH CARD 351</a></td> + <td>351</td> + <td>NIH CARD post-mortem brain (prefrontal cortex); NABEC (European) + HBCC (African/African-admixed), neurologically normal controls</td> + <td>No</td> + <td>~40x ONT (R9.4.1 / R10.4.1)</td> + <td>228,855</td> + <td>1</td> + <td>1</td> + <td>30,282,742</td> +</tr> </table> <p> Note: there is likely some overlap in sample composition across these collections. For example, 1000 Genomes samples are also included in HPRC and CoLoRSdb. </p> <h3><a href="hgTrackUi?g=colorsDbSv">CoLoRSdb SVs</a></h3> <p> Structural variants from the Consortium of Long-Read Sequencing database (CoLoRSdb), from 1,427 PacBio HiFi long-read whole-genome sequences. ~426k SVs (insertions, deletions, inversions) called with pbsv and merged with Jasmine, with allele frequencies, genotype counts and Hardy-Weinberg statistics across the cohort. </p> @@ -334,30 +345,42 @@ <h3><a href="hgTrackUi?g=chirmade101Sv">SVatalog 101 SVs</a></h3> <p> Structural variants from 101 long-read whole-genome sequences released alongside the GWAS SVatalog tool (Chirmade et al. 2026). The samples come from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT), a cystic-fibrosis (CF) patient cohort assembled to model patient-specific responses to CFTR modulator therapies (most participants are F508del homozygotes or F508del / minimal-function compound heterozygotes; a smaller number carry rare nonsense or missense CFTR mutations). ~87k SVs (deletions, insertions, duplications, inversions and complex events) annotated with gene overlaps, ClinGen / gnomAD constraint scores, OMIM / ClinVar / DGV / Decipher regional annotations. </p> +<h3><a href="hgTrackUi?g=cardSv">NIH CARD 351 SVs</a></h3> +<p> +Structural variants from Oxford Nanopore long-read sequencing of post-mortem +brain tissue (prefrontal cortex) from 351 neurologically normal individuals, +generated by the NIH Center for Alzheimer's and Related Dementias (NIH CARD) +Long-Read Initiative (Billingsley et al. 2024). The cohort combines 205 +European-ancestry samples (North American Brain Expression Consortium, NABEC) +and 146 African / African-admixed samples (NIMH Human Brain Collection Core, +HBCC). ~229k SVs (insertions, deletions, inversions) with per-cohort carrier +counts and allele frequencies. +</p> + <h2>Data Access</h2> <p> Each subtrack has its own documentation page with details on how to download and intersect the underlying annotations. The build process for all subtracks is recorded in the UCSC makeDoc, <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/lrSv.txt" target="_blank">doc/hg38/lrSv.txt</a> (and <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hs1/lrSv.txt" target="_blank">doc/hs1/lrSv.txt</a> for T2T-CHM13); the conversion scripts are in <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/lrSv" target="_blank">makeDb/scripts/lrSv</a>, and the track configuration is in <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/lrSv.ra" target="_blank">trackDb/human/lrSv.ra</a>. </p> <h2>References</h2> @@ -482,15 +505,28 @@ </p> <p> Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, Corvelo A, Clarke WE, Musunuri R, Nagulapalli K <em>et al</em>. <a href="https://linkinghub.elsevier.com/retrieve/pii/S0092-8674(22)00991-6" target="_blank"> High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios</a>. <em>Cell</em>. 2022 Sep 1;185(18):3426-3440.e19. PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/36055201" target="_blank">36055201</a>; PMC: <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9439720/" target="_blank">PMC9439720</a> </p> + + +<p> +Billingsley KJ, Meredith M, Daida K, Jerez PA, Negi S, Malik L, Genner RM, Moller A, Zheng X, Gibson +SB <em>et al</em>. +<a href="https://doi.org/10.1101/2024.12.16.628723" target="_blank"> +Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural +variation on gene expression and DNA methylation</a>. +<em>bioRxiv</em>. 2024 Dec 17;. +PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/39764002" target="_blank">39764002</a>; PMC: <a +href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11702628/" target="_blank">PMC11702628</a> +</p> +