0bd565e053abc8c74475f352bedfd37e41312fd2 max Wed Aug 12 02:15:40 2026 -0700 lrSv: update noyvertSv docs and align merged-track source labels, refs #37888 Follow-up to author (Boris Noyvert) feedback on the Noyvert/Boehringer long-read SV dataset. noyvertSv.html: - restore neutral wording about the shared 1000G ONT reads; drop the "independent reprocessing" phrasing and the call-level overlap interpretation the authors objected to - note that singletons (SVs in a single sample) were excluded, so the panel is not exhaustive for the rarest variants - add the medRxiv preprint link alongside the eLife reference Give each dataset one consistent name across its subtrack and the merged (lrSvAll) source filter (databases.tsv + lrSvAll.ra + lrSv.ra): Noyvert 888 (1000G ONT) -> 1KG ONT Boehringer 888 1KG ONT Vienna 1,019 -> 1KG ONT 1019 1KG ONT 100 (Gustafson) -> 1KG ONT UW 100 The gustafsonSv subtrack short/long labels read 97 samples; the paper and our track docs report 100 (Gustafson et al. 2024, PMID 39358015), so those are corrected to 100 as well. Rebuilt lrSvAll.bb with lrSvMergeAll.py; item count unchanged (2,582,278). diff --git src/hg/makeDb/scripts/lrSv/databases.tsv src/hg/makeDb/scripts/lrSv/databases.tsv index a0e381c8c43..f793e5f8acf 100644 --- src/hg/makeDb/scripts/lrSv/databases.tsv +++ src/hg/makeDb/scripts/lrSv/databases.tsv @@ -4,32 +4,32 @@ # label - human-readable label shown in filter dropdown and detail page # bbPath - path to the source bigBed file (under /gbdb/hg38/lrSv/) # valueField - autoSql field name to extract for the per-db value column. # Special values: # UNKNOWN - emit literal "unknown" (db has no real AC) # SPLIT - use affField/healField for two output columns # valueLabel - text shown after the dataset name on the detail page # (typically "AC", but "samples" for gustafson) # affField - SPLIT mode: expression for affected AC (supports "a+b") # healField - SPLIT mode: expression for healthy AC # afField - autoSql field name(s) to use for max-AF aggregation, # comma-separated. "" means no AF available. # Order here = order of per-db columns in the output bigBed. #key label bbPath valueField valueLabel affField healField afField CoLoRSdb CoLoRSdb 1,427 (PacBio) /gbdb/hg38/lrSv/colorsDb/sv.hg38.bb AC AC AF -1000G-ONT-Vienna 1KG ONT Vienna 1,019 /gbdb/hg38/lrSv/1kgOnt.bb AC AC alleleFreq -1000G-ONT 1KG ONT 100 (Gustafson) /gbdb/hg38/lrSv/gustafson.bb sampleCount samples +1000G-ONT-Vienna 1KG ONT 1019 /gbdb/hg38/lrSv/1kgOnt.bb AC AC alleleFreq +1000G-ONT 1KG ONT UW 100 /gbdb/hg38/lrSv/gustafson.bb sampleCount samples AoU1K All of Us 1,027 (PacBio) /gbdb/hg38/lrSv/aou1k.bb AC AC afAfr,afAmr,afEas,afEur,afSas Han945 Han Chinese 945 /gbdb/hg38/lrSv/han945.bb AC AC alleleFreq TommoJapan ToMMo 333 (Japanese) /gbdb/hg38/lrSv/tommoJp.bb AC AC alleleFreq GA4K GA4K 502 (rare disease) /gbdb/hg38/lrSv/ga4kSv.bb AC AC alleleFreq deCODE deCODE 3,622 (Icelandic) /gbdb/hg38/lrSv/decodeSv.bb UNKNOWN AC HPRCv2.1 HPRC v2.1 233 /gbdb/hg38/lrSv/hprc2v21.bb AC AC alleleFreq HGSVC2 HGSVC2 32 /gbdb/hg38/lrSv/hgsvc2.bb AC AC HGSVC3 HGSVC3 65 /gbdb/hg38/lrSv/hgsvc3.bb AC AC # KimPD (Kim PD Brain 100) is held on dev/alpha until published; it has # breakend artifacts up to 190 Mb, so it is excluded from the lrSvAll merge. ArabUAE53 Arab APR 53 /gbdb/hg38/lrSv/apr.bb AC AC alleleFreq China58 CPC 58 (Chinese) /gbdb/hg38/lrSv/cpc1.bb AC AC alleleFreq Svatalog101 SVatalog 101 /gbdb/hg38/lrSv/chirmade101.bb UNKNOWN AC CARD NIH CARD 351 (brain) /gbdb/hg38/lrSv/card.bb AC AC alleleFreq -Noyvert888 Noyvert 888 (1000G ONT) /gbdb/hg38/lrSv/noyvert.bb AC AC AF +Noyvert888 1KG ONT Boehringer 888 /gbdb/hg38/lrSv/noyvert.bb AC AC AF