a8694a3b22d43f0536c02101e9f3b56b5339b4dc
max
  Wed Sep 9 05:47:17 2026 -0700
hubtools: add "import igv" and "splitHap", and the bTaeGut7 zebra finch hub

import igv builds a hub from an IGV session XML. Every Track element becomes a
track, in session order, with the IGV display attributes translated to trackDb
settings. Files the browser can read over the network are linked where they are;
bed, gff, gtf, wig and bedGraph are downloaded and converted, which needs
chrom.sizes and gets them from --chromSizes, from the UCSC assembly, or from a
bigWig of the session itself, the only source there is for a custom assembly.
The BED cleaner exists because real files are not to spec: reversed start/end,
scores over 1000, "#rrggbb" colours, names past 255 characters, and columns that
are not the BED field they sit in, such as trf writing the repeat motif where
thickStart belongs.

splitHap turns a hub built on a diploid assembly into one hub with a genome per
haplotype, reading both assemblies' chrom.sizes and chromAlias from GenArk and
sending each record to whichever assembly has its sequence. It writes
splitHap.report.txt with the records per track per haplotype, the sequences
neither assembly has, and the records reaching past a sequence end, and checks
every track as it goes: records read must equal records matched plus records
with no sequence, and every match must produce an output record or a drop. A
track that does not add up stops the run rather than being written up as a
finding.

Two conversion fixes that came out of the zebra finch data. GFF3 requires unique
IDs, but an annotation of a phased assembly often gives both haplotypes the same
ID; gff3ToGenePred then merges the two copies into one transcript spanning two
chromosomes and discards it, which was losing 31 of 182 retrocopies. IDs that
occur on more than one sequence are now made unique per sequence first. And a
feature name is now taken from the first non-numeric attribute, so a
RepeatMasker GFF gives Motif:Tgut716A rather than the running number in ID=.

genark addContrib gains --tier alpha|beta|public. It edits only betaGenArk.txt
and publicGenArk.txt; beta.hub.txt and public.hub.txt are generated from those
lists and shipped by quickPush.pl, so writing them by hand would push content
outside the normal flow and lose it at the next clade build. The default alpha
tier leaves the lists untouched, so re-running an install cannot demote a
collection that is already promoted.

doc/contrib/bTaeGut7 and trackDb/contrib/bTaeGut7 are the zebra finch
telomere-to-telomere hub built with the above, from the IGV session the authors
ship with the annotations on GenomeArk (Formenti et al, Cell 2026, PMID
42561917). 21 tracks in 6 collections plus 3 standalone, 27 description pages,
and a makeDoc recording where every record went.

diff --git src/hg/makeDb/trackDb/contrib/bTaeGut7/covClr.html src/hg/makeDb/trackDb/contrib/bTaeGut7/covClr.html
new file mode 100644
index 00000000000..52817231ccc
--- /dev/null
+++ src/hg/makeDb/trackDb/contrib/bTaeGut7/covClr.html
@@ -0,0 +1,63 @@
+<h2>Description</h2>
+<p>
+This track is part of the <a href="hgTrackUi?db=$db&amp;g=$parentTrack&amp;hgsid=$hgsid">Read Coverage</a> collection of the zebra finch telomere-to-telomere hub.
+</p>
+<p>
+This track shows coverage of the older PacBio Continuous Long Reads, generated for the
+previous zebra finch reference bTaeGut1.4, after mapping them onto the new assembly. CLR reads
+have a much higher error rate than HiFi reads, and the places where their coverage falls away
+are a direct picture of what the earlier sequencing technology could not reach. Reading it
+next to the Previously Unassembled track shows why those regions were missing.
+</p>
+
+<h2>Display Conventions and Configuration</h2>
+<p>
+A bar graph of read depth, autoscaled to the data in view. Note that these reads come from a
+different individual than the one this genome was assembled from, so some of the variation is
+genuine sequence difference rather than a property of the assembly.
+</p>
+
+<h2>Methods</h2>
+<p>
+PacBio CLR reads generated for the earlier bTaeGut1.4 reference, from a different individual,
+were mapped onto the bTaeGut7 assembly and depth was summarized in fixed windows.
+</p>
+<p>
+The file was taken from GenomeArk at <a href="https://genomeark.s3.amazonaws.com/species/Taeniopygia_guttata/bTaeGut7/manuscript/annotations/bTaeGut1/bTaeGut1.4_CLR.cov.bw" target="_blank">https://genomeark.s3.amazonaws.com/species/Taeniopygia_guttata/bTaeGut7/manuscript/annotations/bTaeGut1/bTaeGut1.4_CLR.cov.bw</a>. It is already in a binary indexed format that the browser reads over the network, so the track points at it where it is and no copy is kept in the hub. The hub is reproduced by the commands in src/hg/makeDb/doc/contrib/bTaeGut7.txt of the UCSC kent source tree.
+</p>
+
+<h2>Data Access</h2>
+<p>
+The data can be explored interactively in table format with the
+<a href="../cgi-bin/hgTables">Table Browser</a> or the
+<a href="../cgi-bin/hgIntegrator">Data Integrator</a> and exported from there to spreadsheet or
+tab-separated tables. From scripts, the data can be accessed through our
+<a href="https://api.genome.ucsc.edu" target="_blank">API</a>, track=<i>covClr</i>.
+</p>
+
+<p>
+This track points straight at the file on GenomeArk rather than at a copy in the hub, so no
+download step is needed to work with it: <tt>bigWigToBedGraph</tt> reads it over the network, as does the browser.
+</p>
+<p>
+The original annotation files are on GenomeArk, <a href="https://genomeark.s3.amazonaws.com/index.html?prefix=species/Taeniopygia_guttata/bTaeGut7/manuscript/annotations/" target="_blank">in the bTaeGut7 annotations directory</a>.
+</p>
+
+<h2>References</h2>
+<p>
+Formenti G, Jain N, Medico JA, Sollitto M, Antipov D, Barcellos S, Biegler M, Borges I, Chang JK,
+Chen Y <em>et al</em>.
+<a href="https://linkinghub.elsevier.com/retrieve/pii/S0092-8674(26)00816-0" target="_blank">
+The complete genome of a songbird</a>.
+<em>Cell</em>. 2026 Aug 6;189(16):4922-4945.e12.
+PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/42561917" target="_blank">42561917</a>
+</p>
+
+<h2>Credits</h2>
+<p>
+The assembly and all of these annotations were produced by the Vertebrate Genome Laboratory at
+The Rockefeller University and their collaborators, and released through GenomeArk. Thanks to
+Giulio Formenti and Erich D. Jarvis and their co-authors for making the data available before
+and after publication. The analysis code is at <a href="https://github.com/gf777/T2T-zebra-finch" target="_blank">github.com/gf777/T2T-zebra-finch</a>. This hub was assembled at UCSC from the IGV
+session distributed with the annotations, using hubtools.
+</p>