73db11cce7b606145e13f1d15b02651f27f58a91 max Wed Sep 16 06:18:38 2026 -0700 danioCode: relabel with full DANIO-CODE branding, un-nest conservation, default on one RNA-seq/ChIP-seq sample Rename the "DC" shortLabel prefix to "DANIO-CODE" throughout (labels were getting long, but the full name matters more than brevity here). Relabel the danioCode superTrack itself to "DANIO-CODE Elements" and drop the now-redundant DANIO-CODE prefix from its remaining nested composites (Elements, Cell Types, COPEs DOPEs, Enhancers, Promoters). Un-nest the Burgess lab's phastCons/CNE conservation data (dcComparativeGenomics) into its own top-level track, "Burgess Fish PhastCons", group compGeno -- same treatment as the CRISPR tracks pulled out earlier, since this isn't DANIO-CODE's own data either. Rewrote every description page's intro: dropped the internal link to the superTrack in favor of a line crediting the DANIO-CODE project and linking to https://danio-code-dcc.genereg.net/, cut each intro to two sentences or fewer, and dropped the general explainer paragraphs (what non-coding regulation is, what conservation is) from danioCode.html and the conservation page entirely. Turn on exactly what the hub itself already marked "on": one RNA-seq sample (Prim-5, Busch-Nentwich lab) and one ChIP-seq signal track (H3K4me3), by letting dcRNAseqComposite and dcChIPseqComposite keep the hub's own "visibility full" instead of forcing hide, via VISIBLE_BY_DEFAULT in danioCodeHubToRa.py. Everything else (hundreds more subtracks per composite) stays off by default. tdbQuery -check passes; still alpha only. refs #38265 diff --git src/hg/makeDb/doc/danRer11/danioCode.txt src/hg/makeDb/doc/danRer11/danioCode.txt index 3d2290b4127..671de9ee787 100644 --- src/hg/makeDb/doc/danRer11/danioCode.txt +++ src/hg/makeDb/doc/danRer11/danioCode.txt @@ -1,187 +1,236 @@ # 2026-09-04 Claude max - DANIO-CODE as native danRer11 tracks (refs #38265, #38239) # The DANIO-CODE consortium publishes its zebrafish developmental genomics data # as a public track hub. This turns the hub's danRer11 trackDb into a native # trackDb file and copies the data files here. # # NOTE: we are hosting the consortium's data, not just pointing at it. #38239 # lists the consortium's agreement to that as an open question; confirm it before # this leaves alpha. mkdir -p /hive/data/genomes/danRer11/bed/danioCode cd /hive/data/genomes/danRer11/bed/danioCode # Get the hub. The server redirects http to https, so use https from here on to # avoid a redirect on every range request the browser makes. curl -sL -O http://trackhub2.genereg.net/DANIO-CODE/DANIO-CODE.hub.txt curl -sL -O http://trackhub2.genereg.net/DANIO-CODE/DANIO-CODE.genomes.txt curl -sL -o trackDb.txt http://trackhub2.genereg.net/DANIO-CODE/danRer11/trackDb.txt grep -cE "^[[:space:]]*track " trackDb.txt # 901 # How much data is behind the hub, and which of it is actually reachable. Pull the # Content-Length with a HEAD request for every bigDataUrl. Takes a few minutes. awk '{gsub(/^ +/,"")} /^bigDataUrl /{u=$2; if (u ~ /^https?:/) print u; else print "https://trackhub2.genereg.net/DANIO-CODE/danRer11/" u}' trackDb.txt \ | sort -u > urls.txt wc -l < urls.txt # 883 cat urls.txt | parallel -j 20 --pipe -N 45 \ 'while read -r u; do sz=$(curl -sIL --max-time 30 "$u" | awk "BEGIN{IGNORECASE=1} /^content-length:/{n=\$2} END{gsub(/\r/,\"\",n); print n+0}") printf "%s\t%s\n" "$sz" "$u" done' > sizes.tsv awk -F'\t' '{s+=$1} END{printf "%.1f GB in %d files\n", s/1e9, NR}' sizes.tsv # 69.3 GB in 883 files # broken down: 702 bigWig 65.3 GB, 1 bw 2.6 GB, 45 bb 1.3 GB, 135 bigBed 0.1 GB # Four files return 404 on the consortium's server. All four are cell-type tracks # whose file name contains a "|" character. Keep the list, the converter drops the # subtracks that point at them. awk -F'\t' '$1==0{print $2}' sizes.tsv > missingUrls.txt cat missingUrls.txt # https://trackhub2.genereg.net/DANIO-CODE/danRer11/tailbud|spinal_cord.colors.bb # https://trackhub2.genereg.net/DANIO-CODE/danRer11/Muscle|tailbud.colors.bb # https://trackhub2.genereg.net/DANIO-CODE/danRer11/Pharyngeal_mesoderm|tailbud.colors.bb # https://trackhub2.genereg.net/DANIO-CODE/danRer11/Epidermis|periderm.colors.bb # Check that the hub's files really are on danRer11 and not lifted from danRer10. bigBedInfo -chroms https://trackhub2.genereg.net/DANIO-CODE/danRer11/consens.canonical.danRer11.bigBed \ | grep -A2 chromCount # chr1 0 59578282 grep -P "^chr1\t" /hive/data/genomes/danRer11/chrom.sizes # chr1 59578282 # matches, no lift needed # The list to download is everything that answered, so 883 minus the 4 above. grep -vxFf missingUrls.txt urls.txt > dlUrls.txt wc -l < dlUrls.txt # 879 # Mirror the data files. 879 files, 69.3 GB, took about an hour at -j 10. The # script skips files that are already there with the right size and prints the # URLs it could not get, so it can be re-run until it reports nothing. ~/kent/src/hg/makeDb/scripts/danioCode/danioCodeDownload.sh \ dlUrls.txt /hive/data/genomes/danRer11/bed/danioCode/data 10 ls data | wc -l # 879 du -sh --apparent-size data # 69G # Check every file against the Content-Length the server reported. python3 -c ' import os want = dict((os.path.basename(u), int(sz)) for sz, u in (l.rstrip("\n").split("\t") for l in open("sizes.tsv"))) bad = [f for f in os.listdir("data") if os.path.getsize("data/" + f) != want.get(f)] print(len(bad), "size mismatches")' # 0 size mismatches # We never parse or rewrite these files, so nothing can shift a coordinate. Spot # check that by re-fetching a few and comparing checksums. for f in copes_dr11.bb transgenic_danRer11.bb consens.canonical.danRer11.bigBed; do curl -s -o /tmp/remote.tmp "https://trackhub2.genereg.net/DANIO-CODE/danRer11/$f" md5sum /tmp/remote.tmp data/$f done # identical for all three # Symlink data/ into /gbdb/danRer11/danioCode/ before the next step, since that is # where the trackDb will point. # Convert the hub trackDb into a native trackDb. What the script has to change is # documented in its header; in short: one superTrack wraps the eleven hub # containers, track names and subGroup tags are made legal for hgTrackDb, relative # bigDataUrls are resolved, settings that tdbQuery rejects for the stanza's type are # dropped, and the subtracks in missingUrls.txt are left out. # --local-prefix points every bigDataUrl at our own copy under /gbdb; leave it off # to keep reading the files from the consortium's server. ~/kent/src/hg/makeDb/scripts/danioCode/danioCodeHubToRa.py \ trackDb.txt \ https://trackhub2.genereg.net/DANIO-CODE/danRer11/ \ danioCode.ra \ --drop-list missingUrls.txt \ --local-prefix /gbdb/danRer11/danioCode # wrote danioCode.ra: 897 stanzas emitted, 4 dropped # 135 subGroup tags renamed # dropped tag 'dimensionXchecked' from 1 stanzas of type 'bigBed 9' # dropped tag 'itemRgb' from 24 stanzas of type 'bigWig' # 901 hub stanzas in, 897 native stanzas out. The four that are missing are the # subtracks whose files 404, listed above. The container superTrack "danioCode" is # added on top, so the trackDb file holds 898 stanzas. cp danioCode.ra ~/kent/src/hg/makeDb/trackDb/zebrafish/danRer11/danioCode.ra # and in zebrafish/danRer11/trackDb.ra: # include danioCode.ra alpha # Two deliberate differences from the hub, both worth revisiting before release: # - all eleven containers are set to "visibility hide", so nothing is on by # default. The hub turns some RNA-seq samples on. Which subset should be # visible is an open question on #38239. # - the three hub superTracks (COPEs/DOPEs, enhancer validation, comparative # genomics) became composites, because a superTrack cannot sit inside another # superTrack. A composite needs its own 'type' -- hgTracks takes the # container's draw handler from it and a hub superTrack carries none -- so the # script borrows the first child's type. Without it hgTracks says # "No draw handler for dcCopes_and_dopes". # # The description page of dcComparativeGenomics has no reference: those tracks are # the Burgess lab's, not DANIO-CODE's, and the citation for them is still unknown. # Check that the browser reads the files and draws them. Worth turning on all # eleven containers at once, which is what caught the missing 'type' above: # https://hgwdev-max.gi.ucsc.edu/cgi-bin/hgRenderTracks?db=danRer11&position=chr1:20,000,000-20,100,000&hideTracks=1&danioCode=show&dcConsensus_promoters=pack&dcComp=full&dcRNAseqComposite=full&DCD002238SQ_pos=full # 2026-09-15 Claude max - un-nest danioCode: give the biggest/most generic # containers their own top-level slot instead of hiding all eleven behind one # superTrack nobody but a DANIO-CODE-aware user would ever open, and split the # Burgess-lab CRISPR tracks out of "DC Conservation" since they aren't # DANIO-CODE's own data. # # RNA-seq -> top-level, group rna # CAGE-seq, 3P-seq -> top-level, group genes # ChIP-seq, Hi-C -> top-level, group regulation # CRISPR (was part of dcComparativeGenomics) -> new top-level superTrack # dcCrispr, group map (mirrors where hg38 keeps its own, unrelated, crispr # tracks), shortLabel/longLabel carry no "DC"/DANIO-CODE branding since # these are the Burgess lab's, not the consortium's # comp, comp_cell_type, copes_and_dopes, evalidation, the remaining # (conservation-only) dcComparativeGenomics, consensus_promoters # -> stay nested under the danioCode superTrack: a mixed # bag of regulatory-element/annotation tracks that # don't map onto one existing group # # Implemented in the converter (STANDALONE_GROUP / NESTED_ORDER / CRISPR_* at # the top of danioCodeHubToRa.py) so it survives the next hub re-import, rather # than hand-editing the generated .ra. Re-run against the cached hub trackDb.txt: cd /hive/data/genomes/danRer11/bed/danioCode ~/kent/src/hg/makeDb/scripts/danioCode/danioCodeHubToRa.py \ trackDb.txt \ https://trackhub2.genereg.net/DANIO-CODE/danRer11/ \ danioCode.ra \ --drop-list missingUrls.txt \ --local-prefix /gbdb/danRer11/danioCode # wrote danioCode.ra: 897 stanzas emitted, 4 dropped (same as before -- the # restructuring only moves stanzas around and adds one synthetic superTrack # stanza, it drops nothing new) cp danioCode.ra ~/kent/src/hg/makeDb/trackDb/zebrafish/danRer11/danioCode.ra # tdbQuery -check requires a superTrack's own children to sit contiguously right # after it in the file -- an unrelated top-level track (one of the newly # standalone composites) sitting between danioCode and one of its remaining # children fails with "X comes between parent (danioCode) and child (Y)". So # TOP_ORDER emits danioCode's own children (NESTED_ORDER) immediately after the # danioCode stanza, then the standalone tracks, then dcCrispr. cd ~/kent/src/hg/makeDb/trackDb tdbQuery -check "select count(*) from danRer11" # no errors; still alpha-only via "include danioCode.ra alpha" in # zebrafish/danRer11/trackDb.ra, so none of this reaches beta/public yet. + +# 2026-09-16 Claude max - relabel and un-nest the Burgess lab conservation track +# +# danioCode superTrack shortLabel -> "DANIO-CODE Elements" (was "DANIO-CODE") +# dcComparativeGenomics (phastCons + CNE, Burgess lab) -> un-nested into its own +# top-level track, shortLabel "Burgess Fish PhastCons", group compGeno +# (Comparative Genomics), alongside dcCrispr as the two Burgess-lab tracks +# distributed via the DANIO-CODE hub but not DANIO-CODE's own data +# the five remaining nested composites (comp, comp_cell_type, copes_and_dopes, +# evalidation, consensus_promoters) drop the "DANIO-CODE " shortLabel prefix, +# since it is now redundant with the superTrack's own label +# +# Implemented in TOP_LABELS / STANDALONE_GROUP / NESTED_ORDER in +# danioCodeHubToRa.py, then regenerated the same way as before: + +cd /hive/data/genomes/danRer11/bed/danioCode +~/kent/src/hg/makeDb/scripts/danioCode/danioCodeHubToRa.py \ + trackDb.txt \ + https://trackhub2.genereg.net/DANIO-CODE/danRer11/ \ + danioCode.ra \ + --drop-list missingUrls.txt \ + --local-prefix /gbdb/danRer11/danioCode +cp danioCode.ra ~/kent/src/hg/makeDb/trackDb/zebrafish/danRer11/danioCode.ra +cd ~/kent/src/hg/makeDb/trackDb +tdbQuery -check "select count(*) from danRer11" +# no errors + +# Also rewrote every description page's intro: dropped the internal link to the +# superTrack in favor of a line crediting the DANIO-CODE project and linking to +# https://danio-code-dcc.genereg.net/, cut each intro to two sentences or fewer, +# and for danioCode.html and dcComparativeGenomics.html (now "Burgess Fish +# PhastCons") dropped the general explainer paragraphs (what non-coding +# regulation is, what conservation is) entirely. + +# Loaded into the personal sandbox to review before committing: +cd ~/kent/src/hg/makeDb/trackDb +make DBS=danRer11 update +# Loaded 2041 track descriptions total; tdbQuery -check clean + +# 2026-09-16 Claude max - turn on one RNA-seq and one ChIP-seq sample by default +# +# Everything was visibility hide (see the note above, still open on #38239 for +# most of the collection). dcRNAseqComposite and dcChIPseqComposite each already +# had exactly one subtrack marked "on" by the hub itself (an RNA-seq sample and +# an H3K4me3 ChIP-seq signal track), so let those two composites keep the hub's +# own "visibility full" instead of forcing hide -- one sample per composite +# shows by default, not the whole pile (RNA-seq alone has 544 subtracks). +# VISIBLE_BY_DEFAULT in danioCodeHubToRa.py controls which composites this +# applies to; regenerated the same way as before and reloaded trackDb_max.