10db3769dd9edfa38ac8d900ca400fdfb514d9a7 gperez2 Wed Jul 29 22:41:46 2026 -0700 An in-place update for gnomad v4.1 to v4.1.1 that swaps bigDataUrl, labels, dataVersion, detailsTabUrls, search descriptions, and removes the alpha-only gnomadVariantsV4.1.1 composite. Updates to gnomad.html, gnomadV4.1.html, and gnomadPLI.html regarding the addition of the gnomad v4.1.1 data, refs #37351 diff --git src/hg/makeDb/trackDb/human/hg38/gnomadPLI.html src/hg/makeDb/trackDb/human/hg38/gnomadPLI.html index 3a3581d274f..48147b6eafe 100644 --- src/hg/makeDb/trackDb/human/hg38/gnomadPLI.html +++ src/hg/makeDb/trackDb/human/hg38/gnomadPLI.html @@ -1,47 +1,47 @@ <H2>Description</H2> <P> The <b>Genome Aggregation Database (gnomAD) - Predicted Constraint Metrics</b> track set contains -metrics of pathogenicity per-gene as predicted for gnomAD v2.1.1, v4.0, or v4.1 and identifies genes subject to -strong selection against various classes of mutation. +metrics of pathogenicity per-gene as predicted for gnomAD v2.1.1, v4.0, v4.1, or v4.1.1 and +identifies genes subject to strong selection against various classes of mutation. </p> <p> This track includes several subtracks of constraint metrics calculated at gene (canonical transcript) and transcript level. For more information see the following <a href="https://macarthurlab.org/2018/10/17/gnomad-v2-1">blog post</a>. The metrics include: <ul> <li>Observed and expected variant counts per transcript/gene <li>Observed/Expected ratio (O/E) <li>Z-scores of the observed counts compared to expected <li>Probability of loss of function intolerance (pLI), for predicted loss-of-function (pLoF) variation only </ul> </p> <h2>Display Conventions and Configuration</h2> <p> -There are two "groups" of tracks in this set, and three gnomAD versions (v2.1.1, v4.0, and v4.1): +There are two "groups" of tracks in this set, and four gnomAD versions (v2.1.1, v4.0, v4.1, and v4.1.1): <ol> <li><b>Gene/Transcript LoF Constraint tracks</b>: Predicted constraint metrics at the whole gene level or whole transcript level for three different types of variation: missense, synonymous, and predicted loss of function. The Gene Constraint track displays metrics for a canonical transcript per gene defined as the longest isoform. The Transcript Constraint track displays metrics for all transcript isoforms. Items on both tracks are shaded according to the pLI score, with outlier items shaded in grey. <br><img src="../images/gnomAD_LOEUFkey.png" width='33%'alt="LOEUF score legend"><br> - Please note there is no gene-level track available for v4.0 and v4.1. + Please note there is no gene-level track available for v4.0, v4.1, or v4.1.1. <li><b>Gene/Transcript Missense Constraint tracks</b>: The missense constraint tracks are built similarly to the LoF constraint tracks, however the items displayed are based on <a href="https://gnomad.broadinstitute.org/help/constraint" target="_blank">missense Z scores</a>. All items are colored black, and individual Z scores can be seen on mouseover. </ol> All tracks follow the general configuration settings for bigBed tracks. Mouseover on the Gene/Transcript Constraint tracks shows the <b>pLI score</b> and the loss of function observed/expected upper bound fraction <b>(LOEUF)</b>, while mouseover on the Regional Constraint track shows only the missense <b>O/E ratio</b>. Clicking on items in any track brings up a table of constraint metrics. </p> <p> Clicking the grey box to the left of the track, or right-clicking and choosing the Configure option, brings up the interface for filtering items based on their pLI score, or labeling the items @@ -112,30 +112,51 @@ to truncating variants. pLI is based on the idea that transcripts can be classified into three categories: <ul> <li>null: heterozygous or homozygous protein truncating variation is completely tolerated <li>recessive: heterozygous variants are tolerated but homozygous variants are not <li>haploinsufficient: heterozygous variants are not tolerated </ul> An expectation-maximization algorithm was then used to assign a probability of belonging in each class to each gene or transcript. pLI is the probability of belonging in the haploinsufficient class. </p> <p> Please see Samocha et al., 2014 and Lek et al., 2016 for further discussion of these metrics. </p> +<h3>Constraint and Gene Quality Flags (v4.1.1)</h3> +<p> +Starting with v4.1.1, the Transcript LoF and Transcript Missense tracks carry two additional flag +fields, shown on the item details page: +<ul> + <li><b>Constraint flags</b>: <code>outlier_lof</code>, <code>outlier_mis</code>, and + <code>outlier_syn</code> mark transcripts whose raw z-score for that class of variation is a + statistical outlier (outside gnomAD's default -5.0 to 5.0 range) and should be interpreted with + caution. <code>no_exp_lof</code>, <code>no_exp_mis</code>, and <code>no_exp_syn</code> mark + transcripts with a missing or zero expected variant count for that class, so constraint metrics + are not calculated for it. + <li><b>Gene flags</b>: <code>low_exome_coverage</code> and <code>low_exome_mapping_quality</code> + mark transcripts overlapping regions of low exome sequencing coverage or low mapping quality. + Constraint metrics for flagged transcripts should be treated with caution. +</ul> +The underlying per-transcript coverage and mapping statistics used to derive these flags +(proportion of bases with allele number at or above 90% of the maximum, mean allele-specific +mapping quality, and proportion of the transcript overlapping segmental duplications or +low-complexity regions) are also shown on the item details page. +</p> + <h3>Transcripts Included</h3> <p> For version 2.1.1 only, the GENCODE transcripts were filtered according to the following criteria: <ul> <li>Must have methionine at start of coding sequence <li>Must have stop codon at end of coding sequence <li>Must be divisible by 3 <li>Must have at least one observed variant when removing exons with median depth < 1 <li>Must have reasonable number of missense and synonymous variants as determined by a Z-score cutoff </ul> </p> <p> For version v2.1.1, the gnomAD gene/transcript data is based on hg19. In order to map transcripts and genes to the hg38 genome the following steps were taken: <ul> <li><b>Transcript track:</b> The gnomAD ENST identifiers were attempted to be matched to all GENCODE versions @@ -183,30 +204,49 @@ </p> <h4>Version based on gnomAD v4.1</h4> <h5>Gene and Transcript Constraint tracks</h5> <p> Per gene and per transcript data were downloaded from the gnomAD Google Storage bucket: <pre> https://storage.googleapis.com/gcp-public-data--gnomad/release/4.1/constraint/gnomad.v4.1.constraint_metrics.tsv </pre> These data were then joined to the Gencode/NCBI set of genes/transcripts available at the UCSC Genome Browser and then transformed into a bigBed 12+5. For the full list of commands used to make this track please see the <a href="https://raw.githubusercontent.com/ucscGenomeBrowser/kent/master/src/hg/makeDb/doc/hg38/gnomad.txt">makedoc</a>. </p> +<h4>Version based on gnomAD v4.1.1</h4> +<h5>Transcript Constraint tracks</h5> +<p> +Per transcript data were downloaded from the gnomAD Google Storage bucket: +<pre> +https://storage.googleapis.com/gcp-public-data--gnomad/release/4.1.1/constraint/gnomad.v4.1.1.constraint_metrics.tsv.bgz +</pre> +gnomAD recomputed the constraint metrics for v4.1.1. The LOEUF calculation changed from a +frequentist approach using a Poisson distribution to a Bayesian framework using a Gamma posterior +distribution. The coverage model was expanded from a depth cutoff to using allele number (AN) as a +proxy for coverage, with a regression model for sites with intermediate coverage. Constraint +metrics are now also available for chrX and chrY. gnomAD's recommended LOEUF threshold changed +from <0.35 to <0.45 accordingly. These data were then joined to the Gencode/NCBI set of +genes/transcripts available at the UCSC Genome Browser and transformed into a bigBed 12+9 +(Transcript LoF) and bigBed 12+8 (Transcript Missense). For the full list of commands used to make +this track please see the +<a href="https://raw.githubusercontent.com/ucscGenomeBrowser/kent/master/src/hg/makeDb/doc/hg38/gnomad.txt">makedoc</a>. +</p> + <h2>Data Access</h2> <p> The raw data can be explored interactively with the <a href="../hgTables">Table Browser</a>, or the <a href="../hgIntegrator">Data Integrator</a>. For automated access, this track, like all others, is available via our <a href="../goldenPath/help/api.html">API</a>. However, for bulk processing, it is recommended to download the dataset. The genome annotation is stored in a bigBed file that can be downloaded from the <a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/gnomAD/pLI/">download server</a>. The exact filenames can be found in the track configuration file. Annotations can be converted to ASCII text by our tool <code>bigBedToBed</code> which can be compiled from the source code or downloaded as a precompiled binary for your system. Instructions for downloading source code and binaries can be found <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">here</a>. The tool can also be used to obtain only features within a given range, for example:</p> <pre> bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gnomAD/pLI/pliByTranscript.bb -chrom=chr6 -start=0 -end=1000000 stdout