442e433a90b25deb87f10e6cf1b7b608bb0a6d67 max Sat Sep 26 21:56:06 2026 -0700 sfariSparkWgs45kAsd: flag 25M insertions of non-human (oral bacteria) sequence as FILTER NonHumanIns and hide them by default; add SFARI SPARK 45k WGS to the combined tracks without those insertions and relabel the 12k pilot as SFARI SPARK iWGS v1.1 Pilot, refs #38424 diff --git src/hg/makeDb/trackDb/human/sfariSparkExomes.html src/hg/makeDb/trackDb/human/sfariSparkExomes.html index d22015a1b5f..0d7df3cd6ba 100644 --- src/hg/makeDb/trackDb/human/sfariSparkExomes.html +++ src/hg/makeDb/trackDb/human/sfariSparkExomes.html @@ -3,31 +3,33 @@ The <a href="https://sparkforautism.org/" target="_blank">Simons Foundation Autism Research Initiative (SFARI)</a> recruited a large cohort of families with autistic children who provided DNA samples and phenotypes. 54,558 families, parents and their children were sequenced, a total of 142,357 individuals with whole-exome (WES) and 12,519 with whole-genome sequencing (WGS). The data contains 32,559 trios and 8,895 quads (one sibling without autism), and 824 twins. </p> <p> The <a href="https://base.sfari.org/dataset/DS0000135" target="_blank">SPARK WGS August 2026 release</a> (SFARI Base dataset DS0000135) adds whole genomes for 45,178 individuals from 21,003 families, 20,858 of them with autism. This is a mostly new set of people: only 14 of them are also in the 12,519-genome release and 201 in the exome release. It includes about 4,900 trios with an autistic child and 1,800 quads. The genomes were sequenced PCR-free on the Illumina NovaSeq X at Broad Clinical Labs. The track "SFARI SPARK 45k WGS ASD" shows this release. Its counts were computed from the genotypes, with the autism/non-autism split -and all variants, including those seen only once (see Methods). +and all variants, including those seen only once. Millions of long insertions in this release +are bacterial DNA from the saliva samples, not human variants. They are marked and hidden by +default (see Methods). </p> <p> The same frequencies shown here are also available publicly on the <a href="https://genomes.sfari.org/" target="_blank">SFARI Genome Browser</a>. See (SPARK et al, Neuron 2018) for details. </p> <h3>Phenotype-stratified counts</h3> <p> In addition to the overall allele count (AC), allele number (AN), and allele frequency (AF), each variant record carries counts split by autism status (the <tt>asd</tt> column of the SPARK individual registration file): </p> <ul> @@ -84,30 +86,45 @@ individuals, with no sample left unassigned. As for the older SPARK tracks, <tt>bcftools +fill-tags</tt> computed AC, AN and AF overall and per group from the called genotypes. Missing genotypes do not count towards AN. The genotypes were then removed. Multiallelic records were split and left-aligned with <tt>bcftools norm</tt>. Alleles that no individual carries after joint genotyping (AC=0) were removed. There is no allele count cutoff. Records that GLnexus marks with the filter <tt>MONOALLELIC</tt> are alleles it could not merge into an overlapping site. In these, only the carriers have a genotype, so AN would count only them and AF would be 1. For these records, AN was set to twice the number of individuals in each group and AF was recomputed from AC. The INFO field <tt>VARLEN</tt> is the length of ALT minus the length of REF: positive for insertions, negative for deletions, 0 for substitutions. The pVCFs were processed in 500 kb pieces on our compute cluster. The scripts are <tt>sparkWgs45kPvcfToSites.sh</tt>, <tt>sparkWgs45kPvcfJobs.sh</tt> and <tt>sparkWgs45kPvcfMerge.sh</tt>. </p> +<p> +This release contains millions of long insertions whose inserted sequence is DNA from +bacteria that live in the mouth, such as <em>Streptococcus mitis</em> and <em>Neisseria +subflava</em>. The DNA was extracted from saliva. We think that reads which are partly +bacterial were soft-clipped by the aligner, and the variant caller then turned the bacterial +part into an insertion. To find them, every inserted sequence of 20 bp or longer was aligned +to the human genome with <tt>bwa mem</tt>. An insertion whose sequence aligns over less than +80% of its length, or with more than 10% mismatches, and which is not also in the gnomAD v4.1 +genomes, gets the value <tt>NonHumanIns</tt> in the FILTER column. These insertions are hidden +by default; to show them, uncheck NonHumanIns under "Exclude variants with these FILTER +values" on this configuration page. The gnomAD check keeps real insertions that are missing +from the reference genome, some of which are common. The script is +<tt>sparkWgs45kFlagNonHumanIns.sh</tt>. +</p> + <p>The sequencing and variant-calling methods are documented as follows by SFARI:</p> <ul> <li> <b>WGS, August 2026 release</b>: DNA was extracted from saliva at Broad Clinical Labs or Prevention Genetics. Genomes were sequenced PCR-free on the Illumina NovaSeq X at Broad Clinical Labs. Samples were kept if contamination was 2.5% or less and 75% or more of reads were mapped. Genetic relationships were checked against the SPARK pedigrees. Reads were realigned to GRCh38 with BWA-MEM2 v2.3, duplicates were marked with GATK MarkDuplicates, and per-sample gVCFs were called with GATK HaplotypeCaller v4.6.2.0. All gVCFs were then joint-called with GLnexus v1.4.1 using its GATK configuration. </li> <li> <b>WGS, 12,519 genomes</b>: This release consists of sequence and variant call data for 12,519 unique individuals, of which 12,517 (99.98%) have available genome-wide