4decf5fbe82051b2d5aff9cabce3feae9dfafae1 max Sat Sep 26 18:48:55 2026 -0700 uniprot: stop declaring the alignments as amino acid coordinates, and show them properly The bigPsl seqType field describes the coordinates, not the letters stored beside them, and the UniProt query side is in bases: these proteins reach the genome through transcripts, so a query runs three bases to a residue. Declaring amino acids made pslFromBigPsl divide the block sizes by three and leave the query coordinates alone, so every reader got an alignment measured in two units at once. That fed a heap overflow in the alignment page, drew blocks short in hgTracks, and loaded sub-codon blocks as size 0, which aborted pslTransMap and took down the lifted SwissProt track (#38249). pslProtFromNaLike() converts such a psl to one counted in residues. Blocks are trimmed to whole codons, and a residue whose codon straddles an exon junction sits in two places in the genome at once, so it gets no column and the page says how many are missing rather than dropping them silently: 0.9% of residues, though 87% of alignments have at least one. A bigPsl also keeps the reference strand where a psl reads the query strand, so minus-strand items arrived claiming the protein was reversed and were rendered as reverse complemented nucleotide ambiguity codes; pslRc moves them to the convention blat uses, query forward and the strand on the target. Checked on hg38 against the translated genome: 210,416 residues over both strands, 99.86% identical, the remainder real protein-vs-reference variation. All 29 items of a test region render the stored protein at the right residues. Files already published still say amino acid and must keep working until they are rebuilt, so the conversion also requires the blocks to measure the target the way the target is measured; verified that separates the two shapes on 3000 records each way, and that all 29 render without crashing in the old format, where they now say plainly that the coordinates and the sequence do not match. refs #38300 diff --git src/inc/psl.h src/inc/psl.h index 57af9d9cfa0..ecc2d07a671 100644 --- src/inc/psl.h +++ src/inc/psl.h @@ -280,30 +280,35 @@ * PSL_CHECK_IGNORE_INSERT_CNTS doesn't validate problems insert counts fields * in each PSL. Useful because protein PSL doesn't seen to compute these in a * consistent way. */ int pslCountBlocks(struct psl *target, struct psl *query, int maxBlockGap); /* count the number of blocks in the query that overlap the target */ /* merge blocks that are closer than maxBlockGap */ struct hash *readPslToBinKeeper(char *sizeFileName, char *pslFileName); /* read a list of psls and return results in hash of binKeeper structure for fast query*/ boolean pslIsProtein(const struct psl *psl); /* is psl a protein psl (are it's blockSizes and scores in protein space) */ +struct psl *pslProtFromNaLike(struct psl *psl, int protSize, int *retDroppedAa); +/* Convert an "na-like" protein psl, whose query is counted in bases at three to a residue, into + * one counted in residues, or NULL if the psl is not that shape. Blocks are trimmed to whole + * codons and the residues that straddle an exon junction are counted in retDroppedAa. */ + struct psl* pslFromAlign(char *qName, int qSize, int qStart, int qEnd, char *qString, char *tName, int tSize, int tStart, int tEnd, char *tString, char* strand, unsigned options); /* Create a PSL from an alignment. Options PSL_IS_SOFTMASK if lower case * bases indicate repeat masking. Returns NULL if alignment is empty after * triming leading and trailing indels.*/ int pslShowAlignment(struct psl *psl, boolean isProt, char *qName, bioSeq *qSeq, int qStart, int qEnd, char *tName, bioSeq *tSeq, int tStart, int tEnd, FILE *f); /* Show protein/DNA alignment or translated DNA alignment in HTML format. */ int pslGenoShowAlignment(struct psl *psl, boolean isProt, char *qName, bioSeq *qSeq, int qStart, int qEnd, char *tName, bioSeq *tSeq, int tStart, int tEnd, int exnStarts[], int exnEnds[], int exnCnt, FILE *f);