f5c96c14557e69252db6935d20ea55bdd250e519 max Fri Sep 4 17:12:57 2026 -0700 DANIO-CODE as native danRer11 tracks, alpha only. Converts the DANIO-CODE consortium's public track hub for danRer11 into a native trackDb: 897 stanzas under one superTrack, with 11 containers for RNA-seq, CAGE-seq, ChIP-seq, 3P-seq, Hi-C, regulatory elements, cell types, COPEs/DOPEs, validated enhancers, conservation and consensus promoters. The 879 data files, 69 GB, are mirrored under /gbdb/danRer11/danioCode and are byte-identical to the consortium's copies. Four cell-type subtracks are left out because their files 404 on the consortium's server. refs #38265 diff --git src/hg/makeDb/trackDb/zebrafish/danRer11/dcCAGEseqComposite.html src/hg/makeDb/trackDb/zebrafish/danRer11/dcCAGEseqComposite.html new file mode 100644 index 00000000000..1efcfa2c301 --- /dev/null +++ src/hg/makeDb/trackDb/zebrafish/danRer11/dcCAGEseqComposite.html @@ -0,0 +1,123 @@ +
+CAGE (cap analysis of gene expression) sequences only the very first bases of capped +RNA molecules. Each read therefore marks one transcription start site, at base +resolution. Because promoters usually fire from a small cluster of neighboring start +sites rather than a single base, the individual start sites are grouped into tag +clusters, and a tag cluster is a good working definition of an active promoter. +
+ ++This track shows CAGE data for 16 zebrafish samples spanning 12 developmental stages, +from the 1-cell stage to the adult. Two kinds of data are shown: the raw signal, which +is the number of transcription start sites seen at each base, and the tag clusters +called from that signal. Zebrafish is a useful system for this because the promoters +used by the mother's stored RNA and those used after the embryo's own genome switches +on can sit within the same promoter region and are separable at this resolution. +
+ ++This track is part of the DANIO-CODE collection. +
+ ++The track has two views that can be configured separately. Signal shows one +coverage graph per sample, auto-scaled to the window. Regions shows the tag +clusters as blocks; the score of a cluster reflects its expression, and the colors are +taken from the consortium's files. +
+ ++Nothing is displayed until samples are selected on the configuration page, where they +can be filtered by developmental stage and by sample. +
+ ++The DANIO-CODE consortium assembled 1,802 zebrafish developmental genomics datasets, +1,438 of them already published and 366 generated by consortium members, and +reprocessed all of them from the raw sequencing reads so that samples from different +laboratories and different protocols can be compared with each other. ChIP-seq and +ATAC-seq were run through the ENCODE pipelines, CAGE-seq through the FANTOM pipeline, +and Hi-C and 4C-seq through the pipelines of the groups that produced them. The +pipelines are published at +gitlab.com/danio-code, and +samples were assigned to developmental stages using ZFIN and ENCODE nomenclature. See +Baranasic et al. 2022 for details. +
+ ++CAGE libraries were processed with the FANTOM CAGE pipeline. Tag clusters were called +from the mapped start sites. The samples in this track come from the Mueller +laboratory and were originally deposited under SRA055273. +
+ ++At UCSC the tracks were converted from the consortium's public track hub at + +trackhub2.genereg.net/DANIO-CODE with the script +danioCodeHubToRa.py, and the data files were copied from the same +server. The data themselves were not modified. The steps are documented in +our makeDoc. +
+ ++The data can be explored interactively in table format with the +Table Browser or the +Data Integrator and exported from there to +spreadsheet or tab-separated tables. From scripts, the data can be accessed through +our API, track=dcCAGEseqComposite. +
+ ++For automated download and analysis, the annotations are stored in bigWig and bigBed files that +can be downloaded from +our +download server. Signal files are named after the DANIO-CODE sample accession, for example DCD001527SQ_signal.bigWig, and tag cluster files end in _tagCluster.bigBed. Individual regions or the whole genome annotation can be +obtained using our tool bigBedToBed, which can be compiled from the source code +or downloaded as a precompiled binary for your system. Instructions for downloading +source code and binaries can be found +here. +The tool can also be used to obtain features within a given range, for example +
+bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/danRer11/danioCode/DCD001527SQ_DCD011313DT_tagCluster.bigBed \ + -chrom=chr1 -start=20000000 -end=20100000 stdout+ +
+The original data files, and the sample and protocol metadata behind them, are +available from the DANIO-CODE data coordination center at +danio-code.zfin.org and from +the consortium's track hub at + +trackhub2.genereg.net/DANIO-CODE. +
+ ++Thanks to the DANIO-CODE consortium for collecting, reprocessing and publishing these +data, and to the laboratories that produced the original datasets. +
+ ++Baranasic D, Hörtenhuber M, Balwierz PJ, Zehnder T, Mukarram AK, Nepal C, Várnai C, Hadzhiev Y, +Jimenez-Gonzalez A, Li N et al. + +Multiomic atlas with functional stratification and developmental dynamics of zebrafish cis- +regulatory elements. +Nat Genet. 2022 Jul;54(7):1037-1050. +PMID: 35789323; PMC: PMC9279159 +
+