1682366b1827b7559f8e1e41635acff6c5ea15e9 max Wed Sep 9 06:05:05 2026 -0700 hprc2annot: move the makeDoc into its own directory and repoint the links The makeDoc has grown a companion (an hg38 pcLAI doc is in progress), so it moves from doc/contrib/hprc2annot.txt into doc/contrib/hprc2annot/, matching how the scripts and trackDb copies are already laid out. The file itself gains a section on the pcLAI scatterplot on the details page: where the reference panel comes from, the four ancestry centroids the discretized field takes across the release, and why the file is read through hgTrackUi rather than fetched by the browser. All seven track description pages linked to the old flat path and would have 404'd, so they are repointed. Six of them change only that link; pclai.html has further edits still in progress and keeps its own copy of the change. refs #35415 diff --git src/hg/makeDb/trackDb/contrib/hprc2annot/liftoffGenes.html src/hg/makeDb/trackDb/contrib/hprc2annot/liftoffGenes.html index 3aaf2f1d15d..b125021e38a 100644 --- src/hg/makeDb/trackDb/contrib/hprc2annot/liftoffGenes.html +++ src/hg/makeDb/trackDb/contrib/hprc2annot/liftoffGenes.html @@ -28,31 +28,31 @@

Methods

Liftoff aligns the transcript sequences of a reference annotation to the target assembly with minimap2 and then chooses, for each gene, the mapping that best preserves its exon-intron structure, optionally identifying additional gene copies. See the reference below for details. For the HPRC pangenome, the human RefSeq annotation (via the CHM13 reference) was lifted onto each assembly.

The annotation files were obtained from the HPRC Release 2 data collection on the public s3://human-pangenomics bucket, indexed at the hprc_intermediate_assembly data tables. Each per-assembly GFF3 was converted to a UCSC bigGenePred file. The Liftoff GFF3 does not record CDS phase, so the phase of each coding exon was recomputed before conversion with gff3ToGenePred and genePredToBigGenePred. The steps are described in the -makeDoc, +makeDoc, the build scripts are in the kent source tree, and the track configuration is in trackDb/contrib/hprc2annot. Roughly one transcript in ten thousand is missing from the track. These are models whose exons Liftoff placed at coordinates that contradict each other, for example a transcript whose recorded start lies after its recorded end, or exons that overlap one another. Such a model cannot be expressed as a gene prediction and is dropped rather than repaired; the counts are recorded per assembly in the build log.

Data Access

For automated analysis, the annotation is stored in a bigBed-format file