e6d1189bea4cc541396f842b65a3392c33c8e734 max Wed Sep 2 02:55:03 2026 -0700 hprc2annot: put the collection in git and fix the QA findings The HPRC Release 2 GenArk contributed track collection (7 tracks x 462 assemblies) had only its one-line betaGenArk.txt enable checked in. Add the makeDoc, the build scripts, the seven track description pages and the trackDb stanzas, and fix the problems QA found. Data fixes, both rebuilt across all 462 assemblies: - liftoff: gff3ToGenePred was naming each genePred after the gene, so every transcript of a gene shared one name, the RefSeq accession was lost and the transcript_biotype lookup never matched (type empty on 99.8% of rows). Pass -rnaNameAttr=ID. Duplicate (chrom,start,end,name) tuples go from 24,969 to 0 and type is now empty on 2,132 of 82,973,730 rows. The same flag is a no-op on the CAT GFF3 (byte-identical output), so both gene tracks now share one code path and CAT needs no rebuild. - segdups: the build read SEDEF column 6, strand1, which is "+" by construction on every row, so every inverted duplication rendered forward. Use column 14, strand2, the orientation of the paralogous copy: 13.8M + and 13.8M - across the collection. Also translate the paralog partner out of PanSN through the GenArk chromAlias, since the browser does not translate a plain text field, and store identity as a percentage so the mouseover can read it. hprc2annotFixBed.sh is not idempotent for pclai: a second run re-parses an already-parsed name and blanks the values. It now refuses to touch a converted file. GCA_041900255.1 was damaged that way and is rebuilt from source. Provenance, all from the QA report: - stats.tsv is appended to rather than truncated on every run, and each run regenerates log/summary.tsv, a per-track roll-up over the collection. - dataVersion on all seven tracks. - Rows are now dropped for exactly two reasons and both are counted: past the end of the sequence, or a sequence name absent from the assembly, which also warns with example names. Only GCA_018472765.3 trips the second, the known upstream contig-version mismatch. genePredToBigGenePred failure is checked and an empty conversion result is a failure, not a valid empty bigBed. Description pages: fix a raw UTF-8 character, rewrite the segdups and pclai display conventions which still described the data before the name field was blanked, add a color legend checked against the data, add the pcLAI preprint (from the Crossref record, since it has no PMID), and correct the stated reason liftoff drops transcripts. Display: title case on the short labels, "Active centromeres" shortened to fit the 17-character limit, pcLAI to pack since it has no readable dense state, liftoff and segdups to dense, and a filter on the segdups original flag. refs #35415 diff --git src/hg/makeDb/trackDb/contrib/hprc2annot/liftoffGenes.html src/hg/makeDb/trackDb/contrib/hprc2annot/liftoffGenes.html new file mode 100644 index 00000000000..3aaf2f1d15d --- /dev/null +++ src/hg/makeDb/trackDb/contrib/hprc2annot/liftoffGenes.html @@ -0,0 +1,83 @@ +<h2>Description</h2> +<p> +This track shows gene annotations mapped onto this Human Pangenome Reference +Consortium (HPRC) Release 2 assembly with Liftoff. Genes are the stretches of DNA +that are copied into RNA and, for most of them, translated into protein. Liftoff +takes an existing, curated set of genes from a reference genome and finds the +matching location of each one in the new assembly, so that familiar RefSeq gene +models can be viewed directly in each individual's genome. Because it maps genes +one by one rather than re-predicting them, Liftoff is well suited to transferring +a trusted reference annotation with its original names and identifiers. +</p> + +<h2>Display Conventions</h2> +<p> +Genes follow the standard UCSC gene display: boxes are exons, connecting lines +are introns, and arrows on the introns show the direction of transcription. +Thicker boxes mark the coding portion (CDS) and thinner boxes the untranslated +regions. When zoomed in, the amino-acid translation and the underlying bases can +be shown. Items are labeled by gene name; the transcript accession (a RefSeq +NM_, NR_, XM_ or XR_ identifier) and the gene and transcript biotypes appear on +the details page. Gene name and transcript accession are both searchable. Liftoff +can place more than one copy of a gene, so some genes appear as several models. +The accession of a transcript sometimes carries a trailing <tt>_1</tt> or similar +suffix, which Liftoff and the source annotation use to keep the identifiers of +such repeated placements distinct. +</p> + +<h2>Methods</h2> +<p> +Liftoff aligns the transcript sequences of a reference annotation to the target +assembly with minimap2 and then chooses, for each gene, the mapping that best +preserves its exon-intron structure, optionally identifying additional gene +copies. See the reference below for details. For the HPRC pangenome, the human +RefSeq annotation (via the CHM13 reference) was lifted onto each assembly. +</p> +<p> +The annotation files were obtained from the HPRC Release 2 data collection on the +public <tt>s3://human-pangenomics</tt> bucket, indexed at +<a href="https://github.com/human-pangenomics/hprc_intermediate_assembly/tree/main/data_tables/annotation/liftoff" target="_blank">the hprc_intermediate_assembly data tables</a>. +Each per-assembly GFF3 was converted to a UCSC bigGenePred file. The Liftoff GFF3 +does not record CDS phase, so the phase of each coding exon was recomputed before +conversion with <tt>gff3ToGenePred</tt> and <tt>genePredToBigGenePred</tt>. The steps are described in the +<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/contrib/hprc2annot.txt" target="_blank">makeDoc</a>, +the build scripts are in the +<a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/hprc2annot" target="_blank">kent source tree</a>, +and the track configuration is in +<a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/trackDb/contrib/hprc2annot" target="_blank">trackDb/contrib/hprc2annot</a>. +Roughly one transcript in ten thousand is missing from the track. These are +models whose exons Liftoff placed at coordinates that contradict each other, for +example a transcript whose recorded start lies after its recorded end, or exons +that overlap one another. Such a model cannot be expressed as a gene prediction +and is dropped rather than repaired; the counts are recorded per assembly in the +build log. +</p> + +<h2>Data Access</h2> +<p> +For automated analysis, the annotation is stored in a bigBed-format file +(<tt>liftoffGenes.bb</tt>) that can be read with the UCSC tool +<tt>bigBedToBed</tt>, which can be compiled from source or downloaded as a +precompiled binary. It can also extract features for a region. The original +annotation files are available from the HPRC S3 bucket linked above. +</p> + +<h2>Credits</h2> +<p> +Annotations were generated by the Human Pangenome Reference Consortium. Thanks to +the HPRC production team for making these data available. +</p> + +<h2>References</h2> + + +<p> +Shumate A, Salzberg SL. +<a href="https://academic.oup.com/bioinformatics/article-lookup/doi/10.1093/bioinformatics/btaa1016" +target="_blank"> +Liftoff: accurate mapping of gene annotations</a>. +<em>Bioinformatics</em>. 2021 Jul 19;37(12):1639-1643. +PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/33320174" target="_blank">33320174</a>; PMC: <a +href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8289374/" target="_blank">PMC8289374</a> +</p> +