360607541994aa88b49fc241c34bd964a1db7b50 lrnassar Tue Sep 29 15:07:07 2026 -0700 Rebuild the hs1 sgdpCopyNumber track as a faceted composite, so its 319 samples are picked from a searchable metadata table rather than 319 checkboxes; the same bigBeds are pointed at by the same /gbdb paths, so no data changed. Subtracks are renamed from ___wssd to sgdpCopyNumber_ because a faceted composite requires the parent name plus the primaryKey value, and dataTypes is deliberately unset: with it hgTrackUi parses the data element only as far as the first underscore and would truncate LP6005441-DNA_A01 to LP6005441-DNA. The per-subtrack 'visibility dense' lines are gone because a faceted composite honors a child's own display mode where a classic composite ignores it, so keeping them would pin every sample to dense and remove the per-item click that the copy number is read from. Sample attributes come from the Reich lab SGDP tables and 317 of the 319 join; sgdpCopyNumberBuild.py takes its sample list from the checked-in sgdpCopyNumberSamples.tsv rather than from trackDb, because at release the generated stanzas replace sgdpCopyNumber.trackDb.ra and the legacy region prefix that the two unmatched samples depend on disappears with them. Alpha gets the new file and beta/public keep the old one until the metadata and color files are on the RR, without which the picker renders empty. sgdpCopyNumber_subset, which turns out to be the first 29 samples in plate order rather than any curated set, is not in the alpha version; whether it is retired for good is still open on the ticket. Faceted composite suggested by Gerardo, and the cross-sandbox metadata fetch that this track turned up was fixed by Max in b3a26a6aff3. refs #29344 diff --git src/hg/makeDb/scripts/sgdpCopyNumber/sgdpCopyNumberFetchMeta.sh src/hg/makeDb/scripts/sgdpCopyNumber/sgdpCopyNumberFetchMeta.sh new file mode 100755 index 00000000000..552026585da --- /dev/null +++ src/hg/makeDb/scripts/sgdpCopyNumber/sgdpCopyNumberFetchMeta.sh @@ -0,0 +1,30 @@ +#!/bin/bash +# Fetch the Simons Genome Diversity Project sample tables used to build the facets +# for the hs1 sgdpCopyNumber track. refs #29344 +# +# The copy-number bigBeds themselves came from the Eichler lab in 2022 and are not +# re-fetched here; only the sample metadata is. +# +# Writes into /hive/data/genomes/hs1/bed/sgdpCopyNumber/metaSrc. + +set -beEu -o pipefail + +destDir=${1:-/hive/data/genomes/hs1/bed/sgdpCopyNumber/metaSrc} +base=https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp + +mkdir -p "$destDir" + +# 344 samples: 279 fully public + 21 signed-letter + 44 from Fan et al. +# Carries Region, Country, Town, Population_ID, Gender, DNA_Source, lat/long, +# keyed on the sequencing library id (Illumina_ID), which is what the bigBed +# file names use. +# +# 280 samples with an ENA/BioSample accession per library id. Used only to add +# a BioSample link to the table; it covers fewer samples than the file above. +for f in SGDP_metadata.279public.21signedLetter.44Fan.samples.txt ena.ftp.pointers.txt ; do + # Download to .part and move into place, so an interrupted fetch can never + # leave a half-written file that looks complete. + curl -sSL --fail -o "$destDir/$f.part" "$base/$f" + mv "$destDir/$f.part" "$destDir/$f" + printf "%s: %d lines\n" "$f" "$(wc -l < "$destDir/$f")" +done