73db11cce7b606145e13f1d15b02651f27f58a91
max
  Wed Sep 16 06:18:38 2026 -0700
danioCode: relabel with full DANIO-CODE branding, un-nest conservation, default on one RNA-seq/ChIP-seq sample

Rename the "DC" shortLabel prefix to "DANIO-CODE" throughout (labels were
getting long, but the full name matters more than brevity here). Relabel the
danioCode superTrack itself to "DANIO-CODE Elements" and drop the now-redundant
DANIO-CODE prefix from its remaining nested composites (Elements, Cell Types,
COPEs DOPEs, Enhancers, Promoters).

Un-nest the Burgess lab's phastCons/CNE conservation data (dcComparativeGenomics)
into its own top-level track, "Burgess Fish PhastCons", group compGeno --
same treatment as the CRISPR tracks pulled out earlier, since this isn't
DANIO-CODE's own data either.

Rewrote every description page's intro: dropped the internal link to the
superTrack in favor of a line crediting the DANIO-CODE project and linking to
https://danio-code-dcc.genereg.net/, cut each intro to two sentences or fewer,
and dropped the general explainer paragraphs (what non-coding regulation is,
what conservation is) from danioCode.html and the conservation page entirely.

Turn on exactly what the hub itself already marked "on": one RNA-seq sample
(Prim-5, Busch-Nentwich lab) and one ChIP-seq signal track (H3K4me3), by
letting dcRNAseqComposite and dcChIPseqComposite keep the hub's own
"visibility full" instead of forcing hide, via VISIBLE_BY_DEFAULT in
danioCodeHubToRa.py. Everything else (hundreds more subtracks per composite)
stays off by default.

tdbQuery -check passes; still alpha only.

refs #38265

diff --git src/hg/makeDb/doc/danRer11/danioCode.txt src/hg/makeDb/doc/danRer11/danioCode.txt
index 3d2290b4127..671de9ee787 100644
--- src/hg/makeDb/doc/danRer11/danioCode.txt
+++ src/hg/makeDb/doc/danRer11/danioCode.txt
@@ -1,187 +1,236 @@
 # 2026-09-04 Claude max - DANIO-CODE as native danRer11 tracks (refs #38265, #38239)
 
 # The DANIO-CODE consortium publishes its zebrafish developmental genomics data
 # as a public track hub.  This turns the hub's danRer11 trackDb into a native
 # trackDb file and copies the data files here.
 #
 # NOTE: we are hosting the consortium's data, not just pointing at it.  #38239
 # lists the consortium's agreement to that as an open question; confirm it before
 # this leaves alpha.
 
 mkdir -p /hive/data/genomes/danRer11/bed/danioCode
 cd /hive/data/genomes/danRer11/bed/danioCode
 
 # Get the hub.  The server redirects http to https, so use https from here on to
 # avoid a redirect on every range request the browser makes.
 curl -sL -O http://trackhub2.genereg.net/DANIO-CODE/DANIO-CODE.hub.txt
 curl -sL -O http://trackhub2.genereg.net/DANIO-CODE/DANIO-CODE.genomes.txt
 curl -sL -o trackDb.txt http://trackhub2.genereg.net/DANIO-CODE/danRer11/trackDb.txt
 
 grep -cE "^[[:space:]]*track " trackDb.txt
 # 901
 
 # How much data is behind the hub, and which of it is actually reachable.  Pull the
 # Content-Length with a HEAD request for every bigDataUrl.  Takes a few minutes.
 awk '{gsub(/^ +/,"")} /^bigDataUrl /{u=$2;
        if (u ~ /^https?:/) print u;
        else print "https://trackhub2.genereg.net/DANIO-CODE/danRer11/" u}' trackDb.txt \
   | sort -u > urls.txt
 wc -l < urls.txt
 # 883
 
 cat urls.txt | parallel -j 20 --pipe -N 45 \
   'while read -r u; do
        sz=$(curl -sIL --max-time 30 "$u" | awk "BEGIN{IGNORECASE=1} /^content-length:/{n=\$2} END{gsub(/\r/,\"\",n); print n+0}")
        printf "%s\t%s\n" "$sz" "$u"
    done' > sizes.tsv
 
 awk -F'\t' '{s+=$1} END{printf "%.1f GB in %d files\n", s/1e9, NR}' sizes.tsv
 # 69.3 GB in 883 files
 # broken down: 702 bigWig 65.3 GB, 1 bw 2.6 GB, 45 bb 1.3 GB, 135 bigBed 0.1 GB
 
 # Four files return 404 on the consortium's server.  All four are cell-type tracks
 # whose file name contains a "|" character.  Keep the list, the converter drops the
 # subtracks that point at them.
 awk -F'\t' '$1==0{print $2}' sizes.tsv > missingUrls.txt
 cat missingUrls.txt
 # https://trackhub2.genereg.net/DANIO-CODE/danRer11/tailbud|spinal_cord.colors.bb
 # https://trackhub2.genereg.net/DANIO-CODE/danRer11/Muscle|tailbud.colors.bb
 # https://trackhub2.genereg.net/DANIO-CODE/danRer11/Pharyngeal_mesoderm|tailbud.colors.bb
 # https://trackhub2.genereg.net/DANIO-CODE/danRer11/Epidermis|periderm.colors.bb
 
 # Check that the hub's files really are on danRer11 and not lifted from danRer10.
 bigBedInfo -chroms https://trackhub2.genereg.net/DANIO-CODE/danRer11/consens.canonical.danRer11.bigBed \
   | grep -A2 chromCount
 #   chr1 0 59578282
 grep -P "^chr1\t" /hive/data/genomes/danRer11/chrom.sizes
 # chr1  59578282
 # matches, no lift needed
 
 # The list to download is everything that answered, so 883 minus the 4 above.
 grep -vxFf missingUrls.txt urls.txt > dlUrls.txt
 wc -l < dlUrls.txt
 # 879
 
 # Mirror the data files.  879 files, 69.3 GB, took about an hour at -j 10.  The
 # script skips files that are already there with the right size and prints the
 # URLs it could not get, so it can be re-run until it reports nothing.
 ~/kent/src/hg/makeDb/scripts/danioCode/danioCodeDownload.sh \
     dlUrls.txt /hive/data/genomes/danRer11/bed/danioCode/data 10
 
 ls data | wc -l
 # 879
 du -sh --apparent-size data
 # 69G
 
 # Check every file against the Content-Length the server reported.
 python3 -c '
 import os
 want = dict((os.path.basename(u), int(sz)) for sz, u in
             (l.rstrip("\n").split("\t") for l in open("sizes.tsv")))
 bad = [f for f in os.listdir("data")
        if os.path.getsize("data/" + f) != want.get(f)]
 print(len(bad), "size mismatches")'
 # 0 size mismatches
 
 # We never parse or rewrite these files, so nothing can shift a coordinate.  Spot
 # check that by re-fetching a few and comparing checksums.
 for f in copes_dr11.bb transgenic_danRer11.bb consens.canonical.danRer11.bigBed; do
     curl -s -o /tmp/remote.tmp "https://trackhub2.genereg.net/DANIO-CODE/danRer11/$f"
     md5sum /tmp/remote.tmp data/$f
 done
 # identical for all three
 
 # Symlink data/ into /gbdb/danRer11/danioCode/ before the next step, since that is
 # where the trackDb will point.
 
 # Convert the hub trackDb into a native trackDb.  What the script has to change is
 # documented in its header; in short: one superTrack wraps the eleven hub
 # containers, track names and subGroup tags are made legal for hgTrackDb, relative
 # bigDataUrls are resolved, settings that tdbQuery rejects for the stanza's type are
 # dropped, and the subtracks in missingUrls.txt are left out.
 # --local-prefix points every bigDataUrl at our own copy under /gbdb; leave it off
 # to keep reading the files from the consortium's server.
 ~/kent/src/hg/makeDb/scripts/danioCode/danioCodeHubToRa.py \
     trackDb.txt \
     https://trackhub2.genereg.net/DANIO-CODE/danRer11/ \
     danioCode.ra \
     --drop-list missingUrls.txt \
     --local-prefix /gbdb/danRer11/danioCode
 # wrote danioCode.ra: 897 stanzas emitted, 4 dropped
 # 135 subGroup tags renamed
 # dropped tag 'dimensionXchecked' from 1 stanzas of type 'bigBed 9'
 # dropped tag 'itemRgb' from 24 stanzas of type 'bigWig'
 
 # 901 hub stanzas in, 897 native stanzas out.  The four that are missing are the
 # subtracks whose files 404, listed above.  The container superTrack "danioCode" is
 # added on top, so the trackDb file holds 898 stanzas.
 
 cp danioCode.ra ~/kent/src/hg/makeDb/trackDb/zebrafish/danRer11/danioCode.ra
 
 # and in zebrafish/danRer11/trackDb.ra:
 #   include danioCode.ra alpha
 
 # Two deliberate differences from the hub, both worth revisiting before release:
 #  - all eleven containers are set to "visibility hide", so nothing is on by
 #    default.  The hub turns some RNA-seq samples on.  Which subset should be
 #    visible is an open question on #38239.
 #  - the three hub superTracks (COPEs/DOPEs, enhancer validation, comparative
 #    genomics) became composites, because a superTrack cannot sit inside another
 #    superTrack.  A composite needs its own 'type' -- hgTracks takes the
 #    container's draw handler from it and a hub superTrack carries none -- so the
 #    script borrows the first child's type.  Without it hgTracks says
 #    "No draw handler for dcCopes_and_dopes".
 #
 # The description page of dcComparativeGenomics has no reference: those tracks are
 # the Burgess lab's, not DANIO-CODE's, and the citation for them is still unknown.
 
 # Check that the browser reads the files and draws them.  Worth turning on all
 # eleven containers at once, which is what caught the missing 'type' above:
 #   https://hgwdev-max.gi.ucsc.edu/cgi-bin/hgRenderTracks?db=danRer11&position=chr1:20,000,000-20,100,000&hideTracks=1&danioCode=show&dcConsensus_promoters=pack&dcComp=full&dcRNAseqComposite=full&DCD002238SQ_pos=full
 
 # 2026-09-15 Claude max - un-nest danioCode: give the biggest/most generic
 # containers their own top-level slot instead of hiding all eleven behind one
 # superTrack nobody but a DANIO-CODE-aware user would ever open, and split the
 # Burgess-lab CRISPR tracks out of "DC Conservation" since they aren't
 # DANIO-CODE's own data.
 #
 #   RNA-seq            -> top-level, group rna
 #   CAGE-seq, 3P-seq    -> top-level, group genes
 #   ChIP-seq, Hi-C      -> top-level, group regulation
 #   CRISPR (was part of dcComparativeGenomics) -> new top-level superTrack
 #     dcCrispr, group map (mirrors where hg38 keeps its own, unrelated, crispr
 #     tracks), shortLabel/longLabel carry no "DC"/DANIO-CODE branding since
 #     these are the Burgess lab's, not the consortium's
 #   comp, comp_cell_type, copes_and_dopes, evalidation, the remaining
 #     (conservation-only) dcComparativeGenomics, consensus_promoters
 #                       -> stay nested under the danioCode superTrack: a mixed
 #                          bag of regulatory-element/annotation tracks that
 #                          don't map onto one existing group
 #
 # Implemented in the converter (STANDALONE_GROUP / NESTED_ORDER / CRISPR_* at
 # the top of danioCodeHubToRa.py) so it survives the next hub re-import, rather
 # than hand-editing the generated .ra.  Re-run against the cached hub trackDb.txt:
 
 cd /hive/data/genomes/danRer11/bed/danioCode
 ~/kent/src/hg/makeDb/scripts/danioCode/danioCodeHubToRa.py \
     trackDb.txt \
     https://trackhub2.genereg.net/DANIO-CODE/danRer11/ \
     danioCode.ra \
     --drop-list missingUrls.txt \
     --local-prefix /gbdb/danRer11/danioCode
 # wrote danioCode.ra: 897 stanzas emitted, 4 dropped (same as before -- the
 # restructuring only moves stanzas around and adds one synthetic superTrack
 # stanza, it drops nothing new)
 
 cp danioCode.ra ~/kent/src/hg/makeDb/trackDb/zebrafish/danRer11/danioCode.ra
 
 # tdbQuery -check requires a superTrack's own children to sit contiguously right
 # after it in the file -- an unrelated top-level track (one of the newly
 # standalone composites) sitting between danioCode and one of its remaining
 # children fails with "X comes between parent (danioCode) and child (Y)".  So
 # TOP_ORDER emits danioCode's own children (NESTED_ORDER) immediately after the
 # danioCode stanza, then the standalone tracks, then dcCrispr.
 cd ~/kent/src/hg/makeDb/trackDb
 tdbQuery -check "select count(*) from danRer11"
 # no errors; still alpha-only via "include danioCode.ra alpha" in
 # zebrafish/danRer11/trackDb.ra, so none of this reaches beta/public yet.
+
+# 2026-09-16 Claude max - relabel and un-nest the Burgess lab conservation track
+#
+#   danioCode superTrack shortLabel -> "DANIO-CODE Elements" (was "DANIO-CODE")
+#   dcComparativeGenomics (phastCons + CNE, Burgess lab) -> un-nested into its own
+#     top-level track, shortLabel "Burgess Fish PhastCons", group compGeno
+#     (Comparative Genomics), alongside dcCrispr as the two Burgess-lab tracks
+#     distributed via the DANIO-CODE hub but not DANIO-CODE's own data
+#   the five remaining nested composites (comp, comp_cell_type, copes_and_dopes,
+#     evalidation, consensus_promoters) drop the "DANIO-CODE " shortLabel prefix,
+#     since it is now redundant with the superTrack's own label
+#
+# Implemented in TOP_LABELS / STANDALONE_GROUP / NESTED_ORDER in
+# danioCodeHubToRa.py, then regenerated the same way as before:
+
+cd /hive/data/genomes/danRer11/bed/danioCode
+~/kent/src/hg/makeDb/scripts/danioCode/danioCodeHubToRa.py \
+    trackDb.txt \
+    https://trackhub2.genereg.net/DANIO-CODE/danRer11/ \
+    danioCode.ra \
+    --drop-list missingUrls.txt \
+    --local-prefix /gbdb/danRer11/danioCode
+cp danioCode.ra ~/kent/src/hg/makeDb/trackDb/zebrafish/danRer11/danioCode.ra
+cd ~/kent/src/hg/makeDb/trackDb
+tdbQuery -check "select count(*) from danRer11"
+# no errors
+
+# Also rewrote every description page's intro: dropped the internal link to the
+# superTrack in favor of a line crediting the DANIO-CODE project and linking to
+# https://danio-code-dcc.genereg.net/, cut each intro to two sentences or fewer,
+# and for danioCode.html and dcComparativeGenomics.html (now "Burgess Fish
+# PhastCons") dropped the general explainer paragraphs (what non-coding
+# regulation is, what conservation is) entirely.
+
+# Loaded into the personal sandbox to review before committing:
+cd ~/kent/src/hg/makeDb/trackDb
+make DBS=danRer11 update
+# Loaded 2041 track descriptions total; tdbQuery -check clean
+
+# 2026-09-16 Claude max - turn on one RNA-seq and one ChIP-seq sample by default
+#
+# Everything was visibility hide (see the note above, still open on #38239 for
+# most of the collection). dcRNAseqComposite and dcChIPseqComposite each already
+# had exactly one subtrack marked "on" by the hub itself (an RNA-seq sample and
+# an H3K4me3 ChIP-seq signal track), so let those two composites keep the hub's
+# own "visibility full" instead of forcing hide -- one sample per composite
+# shows by default, not the whole pile (RNA-seq alone has 544 subtracks).
+# VISIBLE_BY_DEFAULT in danioCodeHubToRa.py controls which composites this
+# applies to; regenerated the same way as before and reloaded trackDb_max.