fa8084416354441c8c817069b7f92344f833d8dc
max
  Tue Sep 8 01:13:43 2026 -0700
hprc2annot: say in the pcLAI docs that uniform color is the expected case

Two of us in a row zoomed a pcLAI track to a few megabases, saw a single flat
color, and concluded the itemRgb was broken. It is not: pclai.bb for
GCA_046629565.1 holds 258 distinct RGB values, but 25,416 of its 25,438 windows
sit in one tight cluster in PCA space (PC1 ~0.40-0.44) and so map to one tight
cluster in color space, R 203-255 G 149-164 B 255, which is a couple of
perceptual steps wide. The color column is a byte-for-byte pass-through of the
HPRC source BED, verified against the S3 file.

Two things the page did not say. A haplotype with one ancestry throughout is
uniform across every chromosome and that is the correct render. And where a
haplotype does carry several ancestries, the blocks are tens of megabases long,
so any view narrower than a chromosome tends to land inside one block and look
uniform too.

Add an example so this is checkable rather than asserted: GCA_018466835.2
(HG02257) has four segment-level ancestry calls, and its chr17 crosses four
blocks with transitions near 10.0, 50.7 and 73.3 Mb. Name the three
chromosomes in that same assembly, chr2 chr13 chr18, that are single-ancestry
end to end, since chr2 is the one that started this.

refs #35415

diff --git src/hg/makeDb/trackDb/contrib/hprc2annot/pclai.html src/hg/makeDb/trackDb/contrib/hprc2annot/pclai.html
index 3d597d26f6a..0b71117dab3 100644
--- src/hg/makeDb/trackDb/contrib/hprc2annot/pclai.html
+++ src/hg/makeDb/trackDb/contrib/hprc2annot/pclai.html
@@ -1,85 +1,106 @@
 <h2>Description</h2>
 <p>
 This track shows pangenome local ancestry inference (pcLAI) for this Human
 Pangenome Reference Consortium (HPRC) Release 2 assembly, in the assembly's own
 coordinates. Every person's genome is a mosaic inherited from ancestors of
 different populations, and "local ancestry" describes, region by region along a
 chromosome, where a stretch of DNA came from. Rather than picking one population
 label per region, pcLAI places each region on a continuous scale, so ancestry
 that sits between two reference populations is not forced into one of them. This
 track divides each haplotype into windows of about a hundred thousand bases and
 reports the ancestry inferred for each one, letting users see how ancestry varies
 along a single individual's genome.
 </p>
 
 <h2>Display Conventions</h2>
 <p>
 Each item is one genomic window. Items carry no visible label. Holding the mouse
 over an item shows the window identifier, the position of the window in the
 principal-component space that pcLAI uses to describe ancestry, the position of
 the longer ancestry segment the window belongs to, and a confidence score; the
 same values are on the details page. The item color is derived from the same
 principal-component position, so windows of similar inferred ancestry get similar
 colors and a run of shared ancestry appears as a block of consistent color.
 Because the color is continuous rather than a set of population labels, it is
 read by comparing regions with each other rather than against a fixed legend.
 </p>
 <p>
+The color follows the ancestry segment rather than the individual window, so
+neighboring windows inside one segment differ by only one or two color steps.
+Those small differences are not meaningful: the feature to read is the block, not
+the window. Two consequences are worth knowing before concluding that a display
+looks wrong. First, a haplotype with a single ancestry throughout gives one
+uniform color across every chromosome, which is the correct result for that
+sample rather than a rendering problem. Second, where a haplotype does carry
+several ancestries the blocks are tens of megabases long, so a view of only a few
+megabases usually falls inside one block and also looks uniform. Zooming out to a
+whole chromosome is what makes the block structure visible.
+</p>
+<p>
+Assembly <tt>GCA_018466835.2</tt> (sample HG02257) is a useful example of a
+haplotype with several ancestry segments. Its <tt>chr17</tt> crosses four blocks,
+with transitions near 10.0 Mb, 50.7 Mb and 73.3 Mb, and <tt>chr5</tt> crosses
+four more across its 184 Mb; <tt>chr20</tt> and <tt>chr22</tt> each cross three.
+In that same assembly <tt>chr2</tt>, <tt>chr13</tt> and <tt>chr18</tt> carry a
+single ancestry from end to end and are uniform in color, so picking one of those
+chromosomes gives no sense of what the track shows.
+</p>
+<p>
 The track is best viewed in pack mode. In dense mode the windows are collapsed
 onto one row, which hides the per-window values.
 </p>
 
 <h2>Methods</h2>
 <p>
 Conventional local ancestry inference gives every segment of a genome one of a
 fixed set of population labels. Point cloud local ancestry inference instead
 places each segment at a point in a continuous coordinate space, so a genome
 becomes a cloud of points, one per haplotype segment. The coordinate space can be
 any continuous description of ancestry; here it is the first two principal
 components of a reference panel of genomes with known population of origin.
 Ancestry that falls between the reference populations, which a label-based method
 has to round to the nearest label, therefore stays visible as an intermediate
 position. The files carry both the coordinate of each individual window and the
 coordinate of the longer ancestry segment that window belongs to. See the
 reference below for the method. This track uses the assembly-coordinate
 (<tt>asm_coord</tt>) output, that is, ancestry placed on each HPRC assembly's own
 sequence; companion outputs projected onto GRCh38 and CHM13 coordinates are
 distributed separately.
 </p>
 <p>
 The annotation files were obtained from the HPRC Release 2 data collection on the
 public <tt>s3://human-pangenomics</tt> bucket, indexed at
 <a href="https://github.com/human-pangenomics/hprc_intermediate_assembly/tree/main/data_tables/annotation/pclai" target="_blank">the hprc_intermediate_assembly data tables</a>.
 The per-assembly BED file was converted to a UCSC bigBed file, with the window
 identifier and the two sets of coordinates split out of the source item name into
 their own fields. The steps are described in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/contrib/hprc2annot.txt" target="_blank">makeDoc</a>,
 the build scripts are in the
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/hprc2annot" target="_blank">kent source tree</a>,
 and the track configuration is in
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/trackDb/contrib/hprc2annot" target="_blank">trackDb/contrib/hprc2annot</a>.
 Every window in the source files is carried through to the bigBed.
 </p>
 
 <h2>Data Access</h2>
 <p>
 For automated analysis, the annotation is stored in a bigBed-format file
 (<tt>pclai.bb</tt>) that can be read with the UCSC tool <tt>bigBedToBed</tt>.
 The original files are available from the HPRC S3 bucket linked above.
 </p>
 
 <h2>Credits</h2>
 <p>
 Annotations were generated by the Human Pangenome Reference Consortium. Thanks to
 the HPRC production team for making these data available.
 </p>
 
 <h2>References</h2>
 <p>
 Geleta M, Mas Montserrat D, Ioannidis NM, Ioannidis AG.
 <a href="https://doi.org/10.64898/2026.03.23.713813" target="_blank">
 Point cloud local ancestry inference (PCLAI): continuous coordinate-based ancestry along the
 genome</a>.
 <em>bioRxiv</em>. 2026 Mar 25.
 doi: <a href="https://doi.org/10.64898/2026.03.23.713813" target="_blank">10.64898/2026.03.23.713813</a>
 </p>