8d5759442b9034d07823c7c57ad4c53f7fc2d6bc max Sat Jul 25 18:25:07 2026 -0700 sfariSparkExomes: update track description diff --git src/hg/makeDb/trackDb/human/sfariSparkExomes.html src/hg/makeDb/trackDb/human/sfariSparkExomes.html index 5a710f64574..1040d98c369 100644 --- src/hg/makeDb/trackDb/human/sfariSparkExomes.html +++ src/hg/makeDb/trackDb/human/sfariSparkExomes.html @@ -39,40 +39,46 @@
Allele frequencies can also be displayed on the SFARI Genome Browser. Full CRAMs and VCFs with genotypes are available from SFARI Base. They require a data access request, which is usually reviewed quickly. More information is available in the SPARK Welcome Packet.
The genome browser track project was approved by the Simons Foundation under request -number 14584.1. WES and WGS data were downloaded from -SFARI Base. -pVCFs were downloaded, anonymized with a script using bcftools and its "fill-tags" plugin and -normalized. There was no minimum allele frequency cutoff. -The ASD-status sample-group file derived from the SPARK individuals_registration -TSV was passed to fill-tags via its -S option, which adds the per-group -AC_AUT/AN_AUT/AF_AUT and AC_NON_AUT/AN_NON_AUT/AF_NON_AUT -tags alongside the overall AC/AN/AF.
+number 14584.1. The multi-sample project VCFs (pVCFs) for both the WES and WGS releases were +downloaded from SFARI Base using Globus. +No minimum allele frequency cutoff was applied. + ++Because the genotype-level pVCFs cannot be redistributed, they were reduced to anonymous, +sites-only VCFs carrying only the overall allele count (AC), allele number (AN) and frequency +(AF), plus the autism-status counts described above, with the bcftools +fill-tags +plugin (its -S option produces the ASD/non-ASD splits), then normalized. The variants +are also annotated with predicted protein consequences using +bcftools csq +against Ensembl gene models; that annotation is displayed in the combined frequency tracks of +this collection. The exact commands for both steps are in the makeDoc file linked below.
-The methods are documented as follows by SFARI:
+The sequencing and variant-calling methods are documented as follows by SFARI:
-The makeDoc file documents how all source files of the varFreqs track were converted. -For some tracks, python scripts were necessary and are also available from GitHub. +The complete, runnable command history for downloading, counting and annotating the SFARI data, +alongside every other cohort in this collection, is in the +makeDoc file; +search it for "SFARI SPARK". The scripts it calls, including +sparkMergeVcfAddCounts.sh (allele counts) and mergeAndAnnotate.sh (merge plus +consequence annotation), are in the +varFreqs scripts directory.
SPARK Consortium. Electronic address: pfeliciano@simonsfoundation.org, SPARK Consortium. SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research. Neuron. 2018 Feb 7;97(3):488-493. PMID: 29420931; PMC: PMC7444276