32048e16a50722894cf28d9b7b4b050302e38c75 markd Wed Sep 30 17:54:58 2026 -0700 Spell the acronym PCLAI, not pcLAI. refs #35415 The authors write it PCLAI throughout: the upstream README at AI-sandbox/hprc-pclai uses PCLAI 18 times and pcLAI never, under the title "Point Cloud Local Ancestry Inference (PCLAI)". We had pcLAI in 517 places, across the track labels, both description pages, the makedocs, the build scripts and the autoSql. hprcPclai.ra holds 467 of those, in shortLabel, longLabel, dataVersion and comments. It is generated, so the fix went into hprcPclaiMakeTrackDb.py and the file was rebuilt; every line of the resulting diff differs only by the rename. Track names, file names and the lowercase pclai in URLs and data file names are untouched. None of them contained the string pcLAI, and renaming the tracks would break saved sessions for no user-visible gain. diff --git src/hg/makeDb/doc/hg38/hprcPclai.txt src/hg/makeDb/doc/hg38/hprcPclai.txt index 4aeca9c4ebc..03e7cd75548 100644 --- src/hg/makeDb/doc/hg38/hprcPclai.txt +++ src/hg/makeDb/doc/hg38/hprcPclai.txt @@ -1,51 +1,51 @@ -# 2026-09-09 Claude max: pcLAI local ancestry for HPRC Release 2 haplotypes, hg38 +# 2026-09-09 Claude max: PCLAI local ancestry for HPRC Release 2 haplotypes, hg38 # One composite track, one subtrack per haplotype. Alpha only. # Experimental / demo: no Redmine ticket, deliberately. If this turns # into a real track the notes below are the whole record of how it was # built and what is known to be off about it. -# The pcLAI (point cloud local ancestry inference) annotations come in three +# The PCLAI (point cloud local ancestry inference) annotations come in three # coordinate systems: each assembly's own coordinates (asm_coord, already shipped # on the GenArk HPRC assembly hubs as the hprc2annot contrib collection), # CHM13 coordinates, and GRCh38 coordinates. This track uses the GRCh38 set. mkdir -p /hive/data/genomes/hg38/bed/hprcPclai cd /hive/data/genomes/hg38/bed/hprcPclai # The index CSV lists one s3:// BED path per haplotype: 231 samples x 2 # haplotypes plus CHM13 = 463 files. Chromosome X and Y are not covered. curl -sL https://raw.githubusercontent.com/human-pangenomics/hprc_intermediate_assembly/refs/heads/main/data_tables/annotation/pclai/pclai_v1.1_grch38_coord_local_hprc_r2.index.csv \ -o pclai_v1.1_grch38_coord_local_hprc_r2.index.csv # Download. The submissions bucket resets connections under load, so the script # retries and runs only 8-wide. 2.5 GB of BED, a couple of minutes. ~/kent/src/hg/makeDb/scripts/hprcPclai/hprcPclaiDownload.sh \ pclai_v1.1_grch38_coord_local_hprc_r2.index.csv bed > download.log 2>&1 grep -c '^OK' download.log # 463 grep FAIL download.log # nothing # Sanity checks before converting: all autosomes, all bed9+1, nothing outside the # chrom bounds, nothing inverted, and the files already sorted. cat bed/*.bed | wc -l # 11936603 cat bed/*.bed | awk '{print NF}' | sort -u # 10 cat bed/*.bed | cut -f1 | sort -u | wc -l # 22, chr1..chr22 awk 'NR==FNR{sz[$1]=$2;next} !($1 in sz){nc++} $3>sz[$1]{over++} $2>=$3{inv++} \ END{print "nochrom",nc+0,"over",over+0,"inverted",inv+0}' \ /hive/data/genomes/hg38/chrom.sizes bed/*.bed # all zero -# The pcLAI BED format is documented at https://github.com/AI-sandbox/hprc-pclai +# The PCLAI BED format is documented at https://github.com/AI-sandbox/hprc-pclai # (README, "Output format"): 0-based half-open like any BED, thickStart == # chromStart, thickEnd == chromEnd, and column 10 is the "centroid" -- the # discretized ancestry of the window, written as the PCA centroid of its ancestry # cluster, which is why it only ever takes four values. Nothing to convert on the # coordinate side. # # thickStart is one base before chromStart in 54811 of the 11936603 windows # (0.46%), against their own spec, and bedToBigBed rejects it. thickEnd always # equals chromEnd, so neither column carries information here; the converter sets # both to the item bounds rather than dropping those windows. cat bed/*.bed | awk -F'\t' '$7!=$2{a++} $7==$2-1{b++} $8!=$3{c++} \ END{print a+0, b+0, c+0}' # 54811 54811 0 # Convert to one bigBed per haplotype. The source name column packs # "SAMPLE/hN/_(PC1,PC2)" and column 10 holds the segment PCA; the script