fa5b31d305066e1938953374d80b3782ef87e239
max
  Mon Sep 7 23:23:40 2026 -0700
Position box: accept a bare codon number, and a range of codon numbers

"KAT6A p.495_533" used to land on codon 495 and silently drop the end of the
range, and a bare codon number after a transcript accession was not understood
at all, so "ENST00000265713.8 p.495" fell through the HGVS code and ended up on
an unrelated locus. Nucleotide ranges already worked. The pseudo-HGVS layer now
takes an optional _end on a bare codon number, and accepts a bare codon number
or range after an NM_ or ENST accession as well as after a gene symbol, looking
up the reference amino acids that HGVS wants and the user did not type.

The accession forms require a literal "p", so "NM_006766.5 1483" keeps meaning
what it meant. A hyphen is still not a range separator: c.1483-1599 is the HGVS
intronic position and stays that way.

Also fixes a read past the end of the protein sequence when the codon number
was larger than the protein, and documents codon ranges in query.html.

refs #38285

diff --git src/hg/lib/tests/input/hgvs/validTerms.txt src/hg/lib/tests/input/hgvs/validTerms.txt
index cbb1de72a45..2c5954c3f80 100644
--- src/hg/lib/tests/input/hgvs/validTerms.txt
+++ src/hg/lib/tests/input/hgvs/validTerms.txt
@@ -17,15 +17,25 @@
 RERE:Thr758Serfs
 GBA1:p.Leu29Alafs*18
 H6PD Val320=
 MTHFR Pro101_Gly102insLeuTyrIleAspValThrTrpHisProAlaGlyAspPro
 YARS1 p.Val153_Val156del
 PLOD1 Glu532del
 PPT1 p.Gln91_Cys96delinsPhe
 PLOD1:p.532
 # From a user via b0b:
 NM_198056.3:pAla735Val
 NM_198056.3:pArg1512Trp
 
 NM_003098.2:c.287G>C
 NM_007294.3:c.2231C>A
 NM_007294.3(BRCA1):c.2231C>A
+
+# Bare codon numbers and ranges of codon numbers, refs #38285
+KAT6A p.495
+KAT6A p.495_533
+KAT6A p.Lys495_Ser533
+NM_006766.5 p.495
+NM_006766.5 p.495_533
+ENST00000265713.8 p.495
+ENST00000265713.8 p.495_533
+KAT6A c.1483_1599