b687dd9018670941ce30f8a4582d6597c5a974d8 lrnassar Tue Sep 1 14:57:43 2026 -0700 lrSv1kLin: fix 2bp insertion span, drop dead numConsolidated field, refresh lrSvAll merge. refs #38099 The Lin 1218 VCFs set INFO/END = POS+1 on insertions, and the converter took chromEnd from END, so every insertion was drawn 2bp wide with svLen 2 instead of the 1bp anchor base. That contradicted the track's own description page and the coordinate convention in the makeDoc, and it kept 107,980 Lin insertions from merging in lrSvAll. Insertions now clamp chromEnd to the anchor; deletions are unchanged and still verify span == |SVLEN| against the source VCFs. Dropped numConsolidated from the converter and the .as: the NumConsolidated INFO key is declared in the VCF header but never appears on a data line, so the column was 0 on all 1.2M rows and added a meaningless line to every detail page. Rebuilt lin1218 on hg38 and hs1 (item counts and variant names unchanged) and re-ran the merge: lrSvAll 2,963,093 -> 2,855,267 rows as the duplicate insertion rows collapse. Bumped seven filter.svLen/insLen maxima in lrSv.ra that were short of the data after the August deletion narrowing, three of them only visible on hs1. lrSvAll.html said the 1000 Genomes linear set was not in the merge, which is no longer true, and gave no warning that sourceCount double-counts because Lin1218 already absorbs HPRC, HGSVC3 and both 1KG ONT callsets. Corrected the merge key description and refreshed ten stale cells in the lrSv.html summary table. diff --git src/hg/makeDb/trackDb/human/lrSv.html src/hg/makeDb/trackDb/human/lrSv.html index 513038ec8bf..813450a9316 100644 --- src/hg/makeDb/trackDb/human/lrSv.html +++ src/hg/makeDb/trackDb/human/lrSv.html @@ -39,31 +39,31 @@ N samples Cohort / disease Disease cases Coverage SV count Min Median Max All merged — All long-read SV datasets merged on identical position+type+length, with per-database AC mixed mixed (HiFi, ONT) - 3,111,026 + 2,855,267 1 1 57,207,413 CoLoRSdb 1,427 Consortium of Long-Read Sequencing, joint callset No mixed (HiFi) 426,239 20 33 101,381 @@ -83,44 +83,44 @@ 1,218 1000 Genomes long-read merge (HiFi assemblies + ONT), Lin et al. No mixed (HiFi assembly, ONT R9/R10) 587,779 50 171 99,968 1KG Vienna ONT 1,019 1000 Genomes, diverse No ~17x ONT - 148,375 + 148,421 2 - 157 - 49,171 + 156 + 49,170 1KG UW ONT 100 1000 Genomes, 5 superpopulations / 19 subpopulations (University of Washington ONT effort) No ~37x ONT (R9.4.1) 113,159 1 - 167 + 166 98,290 1KG Boehringer ONT 888 888 1000 Genomes, 5 superpopulations; used to impute SVs into ~500,000 UK Biobank participants No ~15x ONT (R9.4.1) 107,445 1 1 28,634,664 HPRC v2.1 @@ -162,73 +162,73 @@ No ~40x ONT (R9.4.1 / R10.4.1) 228,855 1 1 30,282,742 deCODE 3,622 3,622 Icelandic general population, no allele counts No ~17x ONT 119,453 1 - 154 - 861,081 + 152 + 861,080 Han 945 945 Han Chinese, general population No ~17x ONT 111,288 1 254 99,744 CPC 58 Chinese Pangenome Consortium, 36 minority ethnic groups (HPRC-specific SVs removed) No ~30x HiFi (pangenome graph) - 36,030 + 36,045 50 - 134 + 130 8,998,096 ToMMo Japanese 333 (111 trios) Japanese, general population No ~22x ONT 74,201 51 158 99,985 Arab APR 53 UAE-resident Arabs from 8 countries (Arab Pangenome Reference) No ~35x HiFi + ~54x ONT (+ Hi-C, pangenome graph) - 72,656 + 72,662 1 121 584,016 GA4K 502 Children's Mercy, pediatric rare disease probands + families Yes (probands) ~27x HiFi 115,554 50 186 809,712