360607541994aa88b49fc241c34bd964a1db7b50
lrnassar
  Tue Sep 29 15:07:07 2026 -0700
Rebuild the hs1 sgdpCopyNumber track as a faceted composite, so its 319
samples are picked from a searchable metadata table rather than 319
checkboxes; the same bigBeds are pointed at by the same /gbdb paths, so no
data changed.  Subtracks are renamed from <region>_<population>_<libId>_wssd
to sgdpCopyNumber_<libId> because a faceted composite requires the parent
name plus the primaryKey value, and dataTypes is deliberately unset: with it
hgTrackUi parses the data element only as far as the first underscore and
would truncate LP6005441-DNA_A01 to LP6005441-DNA.  The per-subtrack
'visibility dense' lines are gone because a faceted composite honors a
child's own display mode where a classic composite ignores it, so keeping
them would pin every sample to dense and remove the per-item click that the
copy number is read from.  Sample attributes come from the Reich lab SGDP
tables and 317 of the 319 join; sgdpCopyNumberBuild.py takes its sample list
from the checked-in sgdpCopyNumberSamples.tsv rather than from trackDb,
because at release the generated stanzas replace sgdpCopyNumber.trackDb.ra
and the legacy region prefix that the two unmatched samples depend on
disappears with them.  Alpha gets the new file and beta/public keep the old
one until the metadata and color files are on the RR, without which the
picker renders empty.  sgdpCopyNumber_subset, which turns out to be the first
29 samples in plate order rather than any curated set, is not in the alpha
version; whether it is retired for good is still open on the ticket.
Faceted composite suggested by Gerardo, and the cross-sandbox metadata fetch
that this track turned up was fixed by Max in b3a26a6aff3.  refs #29344

diff --git src/hg/makeDb/scripts/sgdpCopyNumber/sgdpCopyNumberFetchMeta.sh src/hg/makeDb/scripts/sgdpCopyNumber/sgdpCopyNumberFetchMeta.sh
new file mode 100755
index 00000000000..552026585da
--- /dev/null
+++ src/hg/makeDb/scripts/sgdpCopyNumber/sgdpCopyNumberFetchMeta.sh
@@ -0,0 +1,30 @@
+#!/bin/bash
+# Fetch the Simons Genome Diversity Project sample tables used to build the facets
+# for the hs1 sgdpCopyNumber track.  refs #29344
+#
+# The copy-number bigBeds themselves came from the Eichler lab in 2022 and are not
+# re-fetched here; only the sample metadata is.
+#
+# Writes into /hive/data/genomes/hs1/bed/sgdpCopyNumber/metaSrc.
+
+set -beEu -o pipefail
+
+destDir=${1:-/hive/data/genomes/hs1/bed/sgdpCopyNumber/metaSrc}
+base=https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp
+
+mkdir -p "$destDir"
+
+# 344 samples: 279 fully public + 21 signed-letter + 44 from Fan et al.
+# Carries Region, Country, Town, Population_ID, Gender, DNA_Source, lat/long,
+# keyed on the sequencing library id (Illumina_ID), which is what the bigBed
+# file names use.
+#
+# 280 samples with an ENA/BioSample accession per library id.  Used only to add
+# a BioSample link to the table; it covers fewer samples than the file above.
+for f in SGDP_metadata.279public.21signedLetter.44Fan.samples.txt ena.ftp.pointers.txt ; do
+    # Download to .part and move into place, so an interrupted fetch can never
+    # leave a half-written file that looks complete.
+    curl -sSL --fail -o "$destDir/$f.part" "$base/$f"
+    mv "$destDir/$f.part" "$destDir/$f"
+    printf "%s: %d lines\n" "$f" "$(wc -l < "$destDir/$f")"
+done