95208355e2c667d194b29ee78c8ca8a09c2c2596
max
  Fri Jul 17 08:51:28 2026 -0700
lrSv: add NIH CARD long-read SV subtrack (cardSv)

#Preview2 week - bugs introduced now will need a build patch to fix
Add the NIH CARD Long-Read Initiative structural-variant catalogue (351
post-mortem brain samples: 205 NABEC European ancestry, 146 HBCC
African/African-admixed) as a new subtrack of the Long-read SVs container.
The provider bigBed is re-derived into the shared lrSv schema: signed svLen
made positive (reference span), an explicit insLen added, the single
DUP:TANDEM folded to DUP, and colors remapped to the container's shared
svColor() palette. All 228,855 provider records are carried through 1:1.
Adds the converter and autoSql, the trackDb stanza with filters consistent
with the sibling subtracks, a full description page, a summary row and blurb
on the container page, and a makeDoc section. refs #36258

diff --git src/hg/makeDb/trackDb/human/lrSv.html src/hg/makeDb/trackDb/human/lrSv.html
index 73ecedd397d..615e2c4fd8e 100644
--- src/hg/makeDb/trackDb/human/lrSv.html
+++ src/hg/makeDb/trackDb/human/lrSv.html
@@ -192,30 +192,41 @@
   <td>50</td>
   <td>134</td>
   <td>8,998,096</td>
 </tr>
 <tr>
   <td><a href="hgTrackUi?g=chirmade101Sv">SVatalog 101</a></td>
   <td>101</td>
   <td>Cystic fibrosis (CF) patients from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT). Long-read WGS used for GWAS LD fine-mapping</td>
   <td>Yes (all CF)</td>
   <td>~50x PacBio CLR (34, Sequel I) + ~76x HiFi (67, Sequel II)</td>
   <td>87,068</td>
   <td>4</td>
   <td>160</td>
   <td>1,321,484</td>
 </tr>
+<tr>
+  <td><a href="hgTrackUi?g=cardSv">NIH CARD 351</a></td>
+  <td>351</td>
+  <td>NIH CARD post-mortem brain (prefrontal cortex); NABEC (European) + HBCC (African/African-admixed), neurologically normal controls</td>
+  <td>No</td>
+  <td>~40x ONT (R9.4.1 / R10.4.1)</td>
+  <td>228,855</td>
+  <td>1</td>
+  <td>1</td>
+  <td>30,282,742</td>
+</tr>
 </table>
 
 <p>
 Note: there is likely some overlap in sample composition across these collections.
 For example, 1000 Genomes samples are also included in HPRC and CoLoRSdb.
 </p>
 
 <h3><a href="hgTrackUi?g=colorsDbSv">CoLoRSdb SVs</a></h3>
 <p>
 Structural variants from the Consortium of Long-Read Sequencing database
 (CoLoRSdb), from 1,427 PacBio HiFi long-read whole-genome sequences.
 ~426k SVs (insertions, deletions, inversions) called with pbsv and
 merged with Jasmine, with allele frequencies, genotype counts and
 Hardy-Weinberg statistics across the cohort.
 </p>
@@ -334,30 +345,42 @@
 
 <h3><a href="hgTrackUi?g=chirmade101Sv">SVatalog 101 SVs</a></h3>
 <p>
 Structural variants from 101 long-read whole-genome sequences released
 alongside the GWAS SVatalog tool (Chirmade et al. 2026). The samples come
 from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT), a
 cystic-fibrosis (CF) patient cohort assembled to model patient-specific
 responses to CFTR modulator therapies (most participants are F508del
 homozygotes or F508del / minimal-function compound heterozygotes; a smaller
 number carry rare nonsense or missense CFTR mutations). ~87k SVs
 (deletions, insertions, duplications, inversions and complex events)
 annotated with gene overlaps, ClinGen / gnomAD constraint scores,
 OMIM / ClinVar / DGV / Decipher regional annotations.
 </p>
 
+<h3><a href="hgTrackUi?g=cardSv">NIH CARD 351 SVs</a></h3>
+<p>
+Structural variants from Oxford Nanopore long-read sequencing of post-mortem
+brain tissue (prefrontal cortex) from 351 neurologically normal individuals,
+generated by the NIH Center for Alzheimer's and Related Dementias (NIH CARD)
+Long-Read Initiative (Billingsley et al. 2024). The cohort combines 205
+European-ancestry samples (North American Brain Expression Consortium, NABEC)
+and 146 African / African-admixed samples (NIMH Human Brain Collection Core,
+HBCC). ~229k SVs (insertions, deletions, inversions) with per-cohort carrier
+counts and allele frequencies.
+</p>
+
 
 <h2>Data Access</h2>
 <p>
 Each subtrack has its own documentation page with details on how to download
 and intersect the underlying annotations. The build process for all subtracks
 is recorded in the UCSC makeDoc,
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/lrSv.txt" target="_blank">doc/hg38/lrSv.txt</a>
 (and <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hs1/lrSv.txt" target="_blank">doc/hs1/lrSv.txt</a>
 for T2T-CHM13); the conversion scripts are in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/lrSv" target="_blank">makeDb/scripts/lrSv</a>,
 and the track configuration is in
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/trackDb/human/lrSv.ra" target="_blank">trackDb/human/lrSv.ra</a>.
 </p>
 
 <h2>References</h2>
@@ -482,15 +505,28 @@
 </p>
 
 
 
 <p>
 Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, Corvelo A, Clarke WE, Musunuri R,
 Nagulapalli K <em>et al</em>.
 <a href="https://linkinghub.elsevier.com/retrieve/pii/S0092-8674(22)00991-6" target="_blank">
 High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602
 trios</a>.
 <em>Cell</em>. 2022 Sep 1;185(18):3426-3440.e19.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/36055201" target="_blank">36055201</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9439720/" target="_blank">PMC9439720</a>
 </p>
 
+
+
+<p>
+Billingsley KJ, Meredith M, Daida K, Jerez PA, Negi S, Malik L, Genner RM, Moller A, Zheng X, Gibson
+SB <em>et al</em>.
+<a href="https://doi.org/10.1101/2024.12.16.628723" target="_blank">
+Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural
+variation on gene expression and DNA methylation</a>.
+<em>bioRxiv</em>. 2024 Dec 17;.
+PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/39764002" target="_blank">39764002</a>; PMC: <a
+href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11702628/" target="_blank">PMC11702628</a>
+</p>
+