442e433a90b25deb87f10e6cf1b7b608bb0a6d67 max Sat Sep 26 21:56:06 2026 -0700 sfariSparkWgs45kAsd: flag 25M insertions of non-human (oral bacteria) sequence as FILTER NonHumanIns and hide them by default; add SFARI SPARK 45k WGS to the combined tracks without those insertions and relabel the 12k pilot as SFARI SPARK iWGS v1.1 Pilot, refs #38424 diff --git src/hg/makeDb/trackDb/human/varFreqsAffected.html src/hg/makeDb/trackDb/human/varFreqsAffected.html index c5510bd7279..6a01be6d788 100644 --- src/hg/makeDb/trackDb/human/varFreqsAffected.html +++ src/hg/makeDb/trackDb/human/varFreqsAffected.html @@ -1,33 +1,37 @@

Description

This track shows small variants (SNVs and short indels) that were observed in affected or case individuals of disease-study cohorts, annotated with their predicted protein consequence and colored by severity. It is one half of a matched pair: the companion Population reference track shows the same kind of variants seen in population reference cohorts and in unaffected relatives or controls. Displaying the two together lets you compare, for example, how often a loss-of-function variant in a gene of interest is seen in affected individuals versus the general/unaffected background. For the full list of contributing projects, see the SNV Frequencies collection page.

-The affected counts are drawn from the affected or case arm of five disease-study cohorts: -SFARI SPARK WES and SFARI SPARK WGS (autism spectrum disorder probands), SCHEMA -(schizophrenia cases), GREGoR (affected rare-disease participants), and GA4K (a pediatric -rare-disease cohort). For SPARK, SFARI WGS, SCHEMA, and GREGoR, the source data carries an +The affected counts are drawn from the affected or case arm of six disease-study cohorts: +SFARI SPARK WES, SFARI SPARK iWGS v1.1 Pilot and SFARI SPARK 45k WGS (autism spectrum +disorder probands), SCHEMA (schizophrenia cases), GREGoR (affected rare-disease +participants), and GA4K (a pediatric rare-disease cohort). From SFARI SPARK 45k WGS, the 25 +million insertions flagged as non-human sequence (FILTER NonHumanIns, mostly oral bacteria +DNA from the saliva samples, see the +subtrack's description) were left out. +For the three SPARK cohorts, SCHEMA, and GREGoR, the source data carries an explicit affected/unaffected (or case/control) label, and only the affected arm feeds this track. GA4K reports a single cohort-wide frequency with no per-individual label; because it is a rare-disease cohort, it is counted as affected here, with the caveat that it enrolls parent-child trios, so a minority of its carriers are unaffected parents. Genotyping-array cohorts are not included in either track.

Display Conventions

Color by Consequence

Variants are colored by their most severe predicted consequence:

@@ -49,32 +53,32 @@

Affected AF is the pooled rate across contributing affected arms: affectedAF = sum(AC) / sum(AN), where affectedAC sums the allele counts and affectedAN sums the allele numbers across each cohort/arm that provides both AC and AF (the per-arm AN is derived as round(AC / AF)). Cohorts that publish only AF (with no AC or AN of their own) are still pooled by assigning them an assumed allele number, set as a default_an in the build configuration; their per-arm AC is then derived as round(AF × default_an). Cohorts that publish only AC and have no default_an set (currently GREGoR's per-arm AC_AFFECTED/UNAFFECTED/UNKNOWN) are listed in affectedCohorts but do not contribute to the pool numerator or denominator; their carriers are visible in the per-database AC column instead. The pooled rate is preferred over a max-across-cohorts statistic so a small cohort with a high local AF cannot dominate the displayed frequency.

-The pooled rate also inherits a shared-sample bias: the SFARI SPARK WGS cohort -(~12k probands) is a subset of the larger SFARI SPARK WES cohort (~155k probands), so +The pooled rate also inherits a shared-sample bias: the SFARI SPARK iWGS v1.1 Pilot +cohort (~12k probands) is a subset of the larger SFARI SPARK WES cohort (~155k probands), so probands sequenced in both contribute their AC and AN twice to the affected pool. Where this happens, pooled AN is inflated and pooled AF is skewed toward the frequency in the shared subset. Treat the pooled rate as a cross-cohort summary rather than an unbiased population estimate; the per-cohort AC/AF/AN fields on each variant give the single-cohort numbers.

Top affected sources by AF

Alongside the pooled rate, the mouseover lists the top 3 contributing affected arms ranked by their own per-source AF, formatted as Source (AF). This surfaces case cohorts where the variant is specifically enriched, even when the pooled rate across all arms is small. For disease cohorts that ship a phenotype split (SPARK, SFARI WGS, SCHEMA, GREGoR), the displayed AF is the affected-arm AF and the label includes the

ColorConsequence classExamples
  Protein-truncating / loss-of-function stop_gained, frameshift, splice_donor, splice_acceptor, stop_lost, start_lost