7bef2434fec624473e81cfec3012f5ff0a2c853f gperez2 Mon Sep 14 18:12:47 2026 -0700 Renaming the codon-number tooltip labels to "Genomic codon number" and "Transcript codon number" and rewriting the indel note in plain language, fixing a broken FontAwesome class on the Codon phase link, and rewriting the FAQ's txIndel explanation for accuracy and clarity, refs #38298 diff --git src/hg/htdocs/FAQ/FAQgenes.html src/hg/htdocs/FAQ/FAQgenes.html index 2b84aff651d..d99f1678ad0 100755 --- src/hg/htdocs/FAQ/FAQgenes.html +++ src/hg/htdocs/FAQ/FAQgenes.html @@ -430,79 +430,79 @@ <b>Data format:</b> A small difference is the data format, which matters if you integrate our files into pipelines: The refGene table qName field stores the RefSeq accession but without the version number. The ncbiRefSeq tables show the full accession, with the version number. To add the version number to the refGene table, use a MySQL command like this: <pre> SELECT matches,misMatches,repMatches,nCount,qNumInsert,qBaseInsert,tNumInsert,tBaseInsert,strand,concat(qName, '.', gbSeq.version),qSize,qStart,qEnd,tName,tSize,tStart,tEnd,blockCount,blockSizes,qStarts,tStarts from refSeqAli, hgFixed.gbSeq WHERE refSeqAli.qname=gbSeq.acc</pre> <p>To remove the transcripts on haplotypes, add this condition at the end:</p> <pre>and tName NOT LIKE '%_hap%' AND tName not like '%_alt%' AND tNAME NOT LIKE '%_fix%'</pre> <p>A word of caution on the NCBI RefSeq track on hg19: NCBI is not fully supporting hg19 anymore. As a result, some genes are not located on the main chromosomes anymore. An example is NM_001129826/CSAG3. For hg19, you may prefer UCSC RefSeq for now.</p> <a name="txIndel"></a> <h2>Why does the codon or amino acid number I see differ from the one in a paper or at NCBI?</h2> <p> -Because a RefSeq transcript is a sequence in its own right, not a slice of the genome, the two -can differ in length. Where the transcript carries bases the assembly does not, or the assembly -carries bases the transcript does not, there are two defensible ways to number the codons of -that transcript, and they disagree from the indel onwards: -</p> -<ul> - <li>Counting along the <b>genome</b>, which is what the Genome Browser's gene tracks do. Codons - are counted off the exons as they are laid out on the chromosome, so bases that exist only in - the transcript are never counted. The RefSeq tracks warn you when this happens: the affected - codons are colored orange, an exclamation mark is added after the codon number once you are - zoomed in far enough for the numbers to be drawn, and the codon mouseover gives both - numbers.</li> - <li>Counting along the <b>transcript</b>, which is what NCBI, LOVD, ClinVar and HGVS - <em>c.</em> and <em>p.</em> descriptions do. Codons are counted off the transcript sequence, - including any bases that are missing from the assembly.</li> -</ul> -<p> -Each codon after the difference is shifted by one for every three bases involved. A 21-base -insertion in the transcript, for instance, puts the two counts seven codons apart for the whole -rest of the coding sequence. Neither number is wrong; they answer different questions. But if -you are writing or reading an <b>HGVS</b> description, the transcript count is the one that is -meant, because HGVS <em>c.</em> and <em>p.</em> coordinates are defined on the reference -transcript. +Because a RefSeq transcript is its own independent sequence, not a slice of the genome, the +transcript and the genome can differ in length. Where the transcript carries bases the genome +does not, or the genome +carries bases the transcript does not, there are two valid ways to number the codons of +that transcript: </p> <p> <b>This affects RefSeq only.</b> GENCODE, Ensembl and UCSC Genes transcripts are derived from the -assembly, so their coordinates cannot disagree with it and this problem cannot arise. RefSeq -transcripts are aligned to the assembly instead, so it can, and it affects both the -"NCBI RefSeq" and the "UCSC RefSeq" tracks — often by different amounts, +genome, so their coordinates cannot disagree with it and this problem cannot arise. RefSeq +transcripts are aligned to the genome instead, so it can, and it affects both the +"NCBI RefSeq" and the "UCSC RefSeq" tracks, often by different amounts, because the two use different alignments (see <a href="#ncbiRefseq">the previous question</a>). </p> +<ul> + <li>The <b>genomic codon number</b>, which is what the Genome Browser's gene tracks show by + default. Codons are counted off the exons as they are laid out on the chromosome, so bases + that exist only in the transcript are never counted. The RefSeq tracks warn you when this + happens: the affected codons are colored orange, an exclamation mark is added after the codon + number once you are zoomed in far enough for the numbers to be drawn, and the codon mouseover + gives both numbers.</li> + <li>The <b>transcript codon number</b>, which is what NCBI, LOVD, ClinVar and HGVS + <em>c.</em> and <em>p.</em> descriptions use. Codons are counted off the transcript sequence + itself, including any bases that are missing from the genome.</li> +</ul> +<p> +Before a difference, the two numbers match. At each difference, the gap between them changes, +then stays the same until the next one. A transcript with a single 21-base insertion, for +example, has its two counts a constant seven codons apart for the rest of the coding sequence. A +transcript with several separate differences can have the gap widen more than once along its +length. Neither number is wrong. They answer different questions, but an +HGVS description always means the transcript count, since HGVS <em>c.</em> and <em>p.</em> +coordinates are defined on the reference transcript. +</p> <p> -It is uncommon in human and much more common in other assemblies, and the less complete the -assembly, the more often it happens. In human it clusters in long repetitive coding genes: +It is rare in human, but much more common in other assemblies. The less complete the assembly, +the more often it happens. In human it clusters in long repetitive coding genes: <a href="../cgi-bin/hgTracks?db=hg38&position=chr11:1098960-1099050&refSeqComposite=full" -target="_blank">MUC2</a> carries about 2,900 coding bases that GRCh38 lacks, and FCGBP about +target="_blank">MUC2</a> carries about 2,900 coding bases that hg38 lacks, and FCGBP about 3,500. Shirota <em>et al.</em> catalogued these discrepancies systematically (<a href="https://doi.org/10.1093/database/baw124" target="_blank">Database 2016, baw124</a>, PMID <a href="https://pubmed.ncbi.nlm.nih.gov/27589963/" target="_blank">27589963</a>), and the HGVS nomenclature pages discuss what it means for choosing a <a href="https://hgvs-nomenclature.org/stable/background/refseq/" target="_blank">reference sequence</a>. </p> <p> -The orange marking above lets you tell at a glance that a number you are about to quote is not -the transcript's own. The -underlying insertion or deletion itself is shown in the "RefSeq Alignments" and +The underlying insertion or deletion is itself visible in the "RefSeq Alignments" and "RefSeq Diffs" subtracks of the RefSeq container. Searching the browser for an HGVS term such as <em>NM_002457.5:c.11671A>G</em> uses the transcript numbering and maps it through the alignment, so it lands on the right base either way. </p> <a name="mito"></a> <h2>What is the best gene track for mitochondrial gene annotations</h2> <p> The mitochondrial sequence included in assembly sequence files is a special case and most of what has been explained on this page does not apply to the mitochondrial gene annotations. For most assemblies in the Genome Browser, the sequence name of the mitochondrial genome is "chrM".</p> <p>Both GENCODE and RefSeq databases import their mitochondrial gene annotation directly from the rCRS