059927383e72afe59202535b4863fc016463127a max Thu Sep 3 15:04:05 2026 -0700 Document that cdsStart == cdsEnd marks a non-coding transcript in genePred format, refs #38245 This convention was previously only documented indirectly, as a SQL filtering tip on the Gene tracks FAQ page. Add it next to the cdsStart/cdsEnd field declarations in genePred.as, genePredExt.as, sangerGene.as, ensGene.as, knownGene.as, refFlat.as, genePred.h and sangerGene.h, and mention it in FAQformat.html and bigGenePred.html (via the equivalent thickStart == thickEnd check). diff --git src/hg/htdocs/goldenPath/help/bigGenePred.html src/hg/htdocs/goldenPath/help/bigGenePred.html index 1695e8f97a1..f2f24f9ad6a 100755 --- src/hg/htdocs/goldenPath/help/bigGenePred.html +++ src/hg/htdocs/goldenPath/help/bigGenePred.html @@ -62,30 +62,34 @@ )
The field exonFrames is a comma-separated list of the numbers
with the possible values 0, 1, 2 or -1, one per exon, in order of transcription.
This order means that the first value for a transcript on the minus (-) strand is
the exon on the right of the screen on the Genome Browser.
A value of zero means that the first codon of the exon starts at the first nucleotide of the
exon. A value of one means that the first codon starts after the first
nucleotide and a value of two means that it starts after the second nucleotide.
UTRs are non-coding and their exonFrame value is -1.
The fields cdsStartStat and cdsEndStat have the following values: 'none' = none,
'unk' = unknown, 'incmpl' = incomplete, and 'cmpl' = complete. The
values, however, are not used for our display and cannot be used to identify coding or non-coding genes.
+To determine whether a transcript is non-coding, check whether thickStart equals
+thickEnd (coding transcripts have thickStart != thickEnd); this is the
+bigGenePred equivalent of the genePred convention where cdsStart == cdsEnd marks a
+non-coding transcript.
For most purposes, to get more information about a transcript, other tables will need to be used. For
instance, in the case of hg38, the tables named wgEncodeGencodeAttrsVxx, where xx is the Gencode Version number.
See this coding/non-coding genes FAQ
for more information.
The following bed12+8 is an example of a pre-bigGenePred text file .
Step 1.
Format your pre-bigGenePred file. The first 12 fields of pre-bigGenePred files are described by the
BED file format. Your file must
also contain the 8 extra fields described in the autoSql file definition
shown above: name2, cdsStartStat, cdsEndStat, exonFrames, type, geneName, geneName2,