7bef2434fec624473e81cfec3012f5ff0a2c853f gperez2 Mon Sep 14 18:12:47 2026 -0700 Renaming the codon-number tooltip labels to "Genomic codon number" and "Transcript codon number" and rewriting the indel note in plain language, fixing a broken FontAwesome class on the Codon phase link, and rewriting the FAQ's txIndel explanation for accuracy and clarity, refs #38298 diff --git src/hg/htdocs/FAQ/FAQgenes.html src/hg/htdocs/FAQ/FAQgenes.html index 2b84aff651d..d99f1678ad0 100755 --- src/hg/htdocs/FAQ/FAQgenes.html +++ src/hg/htdocs/FAQ/FAQgenes.html @@ -430,79 +430,79 @@ Data format: A small difference is the data format, which matters if you integrate our files into pipelines: The refGene table qName field stores the RefSeq accession but without the version number. The ncbiRefSeq tables show the full accession, with the version number. To add the version number to the refGene table, use a MySQL command like this:
SELECT matches,misMatches,repMatches,nCount,qNumInsert,qBaseInsert,tNumInsert,tBaseInsert,strand,concat(qName, '.', gbSeq.version),qSize,qStart,qEnd,tName,tSize,tStart,tEnd,blockCount,blockSizes,qStarts,tStarts from refSeqAli, hgFixed.gbSeq WHERE refSeqAli.qname=gbSeq.acc
To remove the transcripts on haplotypes, add this condition at the end:
and tName NOT LIKE '%_hap%' AND tName not like '%_alt%' AND tNAME NOT LIKE '%_fix%'
A word of caution on the NCBI RefSeq track on hg19: NCBI is not fully supporting hg19 anymore. As a result, some genes are not located on the main chromosomes anymore. An example is NM_001129826/CSAG3. For hg19, you may prefer UCSC RefSeq for now.
-Because a RefSeq transcript is a sequence in its own right, not a slice of the genome, the two -can differ in length. Where the transcript carries bases the assembly does not, or the assembly -carries bases the transcript does not, there are two defensible ways to number the codons of -that transcript, and they disagree from the indel onwards: -
--Each codon after the difference is shifted by one for every three bases involved. A 21-base -insertion in the transcript, for instance, puts the two counts seven codons apart for the whole -rest of the coding sequence. Neither number is wrong; they answer different questions. But if -you are writing or reading an HGVS description, the transcript count is the one that is -meant, because HGVS c. and p. coordinates are defined on the reference -transcript. +Because a RefSeq transcript is its own independent sequence, not a slice of the genome, the +transcript and the genome can differ in length. Where the transcript carries bases the genome +does not, or the genome +carries bases the transcript does not, there are two valid ways to number the codons of +that transcript:
This affects RefSeq only. GENCODE, Ensembl and UCSC Genes transcripts are derived from the -assembly, so their coordinates cannot disagree with it and this problem cannot arise. RefSeq -transcripts are aligned to the assembly instead, so it can, and it affects both the -"NCBI RefSeq" and the "UCSC RefSeq" tracks — often by different amounts, +genome, so their coordinates cannot disagree with it and this problem cannot arise. RefSeq +transcripts are aligned to the genome instead, so it can, and it affects both the +"NCBI RefSeq" and the "UCSC RefSeq" tracks, often by different amounts, because the two use different alignments (see the previous question).
++Before a difference, the two numbers match. At each difference, the gap between them changes, +then stays the same until the next one. A transcript with a single 21-base insertion, for +example, has its two counts a constant seven codons apart for the rest of the coding sequence. A +transcript with several separate differences can have the gap widen more than once along its +length. Neither number is wrong. They answer different questions, but an +HGVS description always means the transcript count, since HGVS c. and p. +coordinates are defined on the reference transcript. +
-It is uncommon in human and much more common in other assemblies, and the less complete the -assembly, the more often it happens. In human it clusters in long repetitive coding genes: +It is rare in human, but much more common in other assemblies. The less complete the assembly, +the more often it happens. In human it clusters in long repetitive coding genes: MUC2 carries about 2,900 coding bases that GRCh38 lacks, and FCGBP about +target="_blank">MUC2 carries about 2,900 coding bases that hg38 lacks, and FCGBP about 3,500. Shirota et al. catalogued these discrepancies systematically (Database 2016, baw124, PMID 27589963), and the HGVS nomenclature pages discuss what it means for choosing a reference sequence.
-The orange marking above lets you tell at a glance that a number you are about to quote is not -the transcript's own. The -underlying insertion or deletion itself is shown in the "RefSeq Alignments" and +The underlying insertion or deletion is itself visible in the "RefSeq Alignments" and "RefSeq Diffs" subtracks of the RefSeq container. Searching the browser for an HGVS term such as NM_002457.5:c.11671A>G uses the transcript numbering and maps it through the alignment, so it lands on the right base either way.
The mitochondrial sequence included in assembly sequence files is a special case and most of what has been explained on this page does not apply to the mitochondrial gene annotations. For most assemblies in the Genome Browser, the sequence name of the mitochondrial genome is "chrM".
Both GENCODE and RefSeq databases import their mitochondrial gene annotation directly from the rCRS