7aa59c8f6afbda4c2157d3990b2bbc48a39cac63
max
  Fri Sep 25 02:48:31 2026 -0700
varFreqs: add SFARI SPARK 45k WGS subtracks. sfariSparkWgs45k is built from the release's AF table (AN estimated, singletons dropped); sfariSparkWgs45kAsd is built from the genotype pVCFs with ASD/non-ASD counts, for now only the DSCAM locus while the genome-wide parasol run finishes, refs #38424

diff --git src/hg/makeDb/trackDb/human/sfariSparkExomes.html src/hg/makeDb/trackDb/human/sfariSparkExomes.html
index 1040d98c369..b833859acbe 100644
--- src/hg/makeDb/trackDb/human/sfariSparkExomes.html
+++ src/hg/makeDb/trackDb/human/sfariSparkExomes.html
@@ -1,153 +1,209 @@
 <h2>Description</h2>
 <p>
 The <a href="https://sparkforautism.org/" target="_blank">Simons Foundation Autism Research
 Initiative (SFARI)</a> recruited a large cohort of families with autistic children who provided
 DNA samples and phenotypes. 54,558 families, parents and their children were sequenced, a total
 of 142,357 individuals with whole-exome (WES) and 12,519 with whole-genome sequencing (WGS).
 The data contains 32,559 trios and 8,895 quads (one sibling without autism), and 824 twins.
 </p>
 
+<p>
+The <a href="https://base.sfari.org/dataset/DS0000135" target="_blank">SPARK WGS August 2026
+release</a> (SFARI Base dataset DS0000135) adds whole genomes for 45,178 individuals from
+21,003 families, 20,858 of them with autism. This is a mostly new set of people: only 14 of
+them are also in the 12,519-genome release and 201 in the exome release. It includes about
+4,900 trios with an autistic child and 1,800 quads. The genomes were sequenced PCR-free on
+the Illumina NovaSeq X at Broad Clinical Labs. Two tracks show this release. "SFARI SPARK 45k
+WGS" was made from a frequency table that SFARI released with the data. It has no
+autism/non-autism split, its allele counts are estimates, and variants seen only once are
+left out. "SFARI SPARK 45k WGS ASD" was computed from the genotypes. It has exact counts,
+the autism/non-autism split and all variants, including those seen once. For now it covers
+only the DSCAM gene on chr21 (see Methods).
+</p>
+
 <p>
 The same frequencies shown here are also available publicly on the
 <a href="https://genomes.sfari.org/" target="_blank">SFARI Genome Browser</a>.
 See (SPARK et al, Neuron 2018) for details.
 </p>
 
 <h3>Phenotype-stratified counts</h3>
 <p>
 In addition to the overall allele count (AC), allele number (AN), and allele
 frequency (AF), each variant record carries counts split by autism status
 (the <tt>asd</tt> column of the SPARK individual registration file):
 </p>
 <ul>
   <li><tt>AC_AUT</tt>, <tt>AN_AUT</tt>, <tt>AF_AUT</tt> &mdash; individuals
       coded as autistic (<tt>asd=TRUE</tt>).</li>
   <li><tt>AC_NON_AUT</tt>, <tt>AN_NON_AUT</tt>, <tt>AF_NON_AUT</tt> &mdash;
       individuals coded as non-autistic (<tt>asd=FALSE</tt>); these are
       mostly parents and unaffected siblings of the probands.</li>
 </ul>
 <p>
 A small minority of samples have a blank <tt>asd</tt> value and so contribute
 only to the overall AC/AN/AF, not to either group total.
 </p>
 
 <h2>Data Access</h2>
 <p>
 Due to license restrictions, the data for this track cannot be downloaded from the UCSC
 Genome Browser. The Table Browser, Data Integrator, and download server are not available
 for this track.
 </p>
 <p>
 Allele frequencies can also be displayed on the
 <a href="https://genomes.sfari.org/" target="_blank">SFARI Genome Browser</a>.
 Full CRAMs and VCFs with genotypes are available from
 <a href="https://base.sfari.org/" target="_blank">SFARI Base</a>.
 They require a data access request, which is usually reviewed quickly. More information is
 available in the
 <a href="https://cohorts-cdn.simonsfoundation.org/spark/researcher_packets/SPARK_SFARI_Researcher_Welcome_Packet.pdf"
 target="_blank">SPARK Welcome Packet</a>.
 </p>
 
 <h2>Methods</h2>
 
 <p>The genome browser track project was approved by the Simons Foundation under request
-number 14584.1. The multi-sample project VCFs (pVCFs) for both the WES and WGS releases were
+number 14584.1. The multi-sample project VCFs (pVCFs) for the 2024 WES and the 12,519-genome WGS releases were
 downloaded from <a href="https://base.sfari.org/" target="_blank">SFARI Base</a> using Globus.
 No minimum allele frequency cutoff was applied.</p>
 
 <p>
 Because the genotype-level pVCFs cannot be redistributed, they were reduced to anonymous,
 sites-only VCFs carrying only the overall allele count (AC), allele number (AN) and frequency
 (AF), plus the autism-status counts described above, with the bcftools
 <a href="http://samtools.github.io/bcftools/howtos/plugin.fill-tags.html" target="_blank">fill-tags</a>
 plugin (its <tt>-S</tt> option produces the ASD/non-ASD splits), then normalized. The variants
 are also annotated with predicted protein consequences using
 <a href="http://samtools.github.io/bcftools/bcftools.html#csq" target="_blank">bcftools csq</a>
 against Ensembl gene models; that annotation is displayed in the combined frequency tracks of
 this collection. The exact commands for both steps are in the makeDoc file linked below.</p>
 
+<p>
+The WGS August 2026 release is handled differently. It was downloaded from SFARI Base with
+Globus, but only the table <tt>SPARK.WGS.2026_08.gatk.pvcf_variant_frequencies.tsv</tt> was
+used. SFARI made this table from the pVCFs with <tt>bcftools query</tt>. It lists chromosome,
+position, REF, ALT and AF, with AF rounded to six decimals. It has no allele count or allele
+number, so we estimated them. At every site, including sites on chrX and chrY, all low
+frequencies are exact multiples of 1/90,356, and 90,356 is two alleles for each of the
+45,178 individuals. For example, all of the 179 million alleles seen once have AF=1.1e-05.
+None has 1.2e-05, which is what a site with more than about 4% missing genotypes would give.
+We therefore set AN=90,356 for every variant and computed AC as AF &times; AN, rounded to the
+nearest integer. Multiallelic records were split into one record per alternate allele.
+Alleles with an estimated AC of 1 were removed. The remaining alleles were left-aligned
+against the reference with <tt>bcftools norm</tt>. The script is
+<tt>sparkWgs45kToVcf.sh</tt>.
+</p>
+
+<p>
+The track "SFARI SPARK 45k WGS ASD" was made from the genotype-level pVCFs of the same
+release, which are split into 2.5 Mb chunks. The samples were grouped by the <tt>asd</tt>
+column of the release's sample metadata file: 20,858 autistic and 24,320 non-autistic
+individuals, with no sample left unassigned. As for the older SPARK tracks, <tt>bcftools
++fill-tags</tt> computed AC, AN and AF overall and per group from the called genotypes.
+Missing genotypes do not count towards AN. The genotypes were then removed. Multiallelic
+records were split and left-aligned with <tt>bcftools norm</tt>. Alleles that no individual
+carries after joint genotyping (AC=0) were removed. There is no allele count cutoff. The
+INFO field <tt>VARLEN</tt> is the length of ALT minus the length of REF: positive for
+insertions, negative for deletions, 0 for substitutions. Unlike the other SPARK tracks,
+this track also shows the predicted protein consequences (<tt>BCSQ</tt>) on its details
+page, computed with <tt>bcftools csq</tt> against Ensembl release 115. The scripts are
+<tt>sparkWgs45kPvcfToSites.sh</tt> and <tt>sparkWgs45kPvcfRange.sh</tt>.
+</p>
+
 <p>The sequencing and variant-calling methods are documented as follows by SFARI:</p>
 <ul>
   <li>
-    <b>WGS</b>:
+    <b>WGS, August 2026 release</b>: DNA was extracted from saliva at Broad Clinical Labs or
+    Prevention Genetics. Genomes were sequenced PCR-free on the Illumina NovaSeq X at Broad
+    Clinical Labs. Samples were kept if contamination was 2.5% or less and 75% or more of
+    reads were mapped. Genetic relationships were checked against the SPARK pedigrees. Reads
+    were realigned to GRCh38 with BWA-MEM2 v2.3, duplicates were marked with GATK
+    MarkDuplicates, and per-sample gVCFs were called with GATK HaplotypeCaller v4.6.2.0. All
+    gVCFs were then joint-called with GLnexus v1.4.1 using its GATK configuration.
+  </li>
+  <li>
+    <b>WGS, 12,519 genomes</b>:
     This release consists of sequence and variant call data for 12,519
     unique individuals, of which 12,517 (99.98%) have available genome-wide
     SNP genotype data. Sequencing and genotyping of all samples in this
     release was performed at New York Genome Center (NYGC). DNA from saliva
     samples were extracted and prepared with PCR-free methods and sequenced
     with paired-end sequencing of 150 bases on the Illumina NovaSeq 6000
     system. Alignment of reads to the human reference genome version
     GRCh38, duplicate read marking, and Base Quality Score Recalibration
     (BQSR) were performed by New York Genome Center (NYGC). Whole-genome
     sequencing data were processed using a standardized, functionally
     equivalent CCDG pipeline with alignment to the GRCh38DH (1000 Genomes)
     reference using BWA-MEM v0.7.15 (deterministic settings, no -M, use of
     .alt contigs), Picard-equivalent duplicate marking (Picard &ge;2.4.1 or
     equivalent), no indel realignment, and base quality score recalibration
     with GATK (dbSNP138, Mills and 1000G gold-standard indels, known
     indels). Final outputs were stored as lossless CRAM files with
     complete SAM-compliant read-group annotations and mandatory 4-bin
     base-quality compression (Q2&mdash;6, 10, 20, 30), and all implementations
     were validated for functional equivalence across centers before use.
     Variant Calling was performed using DeepVariant. See
     <a href="https://github.com/CCDG/Pipeline-Standardization/blob/master/PipelineStandard.md"
     target="_blank">CCDG pipeline details</a>.
   </li>
   <li>
     <b>WES</b>: This release contains
     sequence data for 142,357 individuals and genotyping data for
     141,368 individuals. DNA was sequenced from saliva for all
     samples and all participants consented to having their genetic
     data shared by Regeneron. Exomes for all samples were sequenced with
     short-read, paired-end sequencing of 150 bases on Illumina
     NovaSeq 6000 machines using S2/S4 flow cells. Sequencing and
     genotyping was performed across nine batches (WES1 through
     WES9) at the Regeneron Genetics Center (RGC) and integrated
     together for this data release. All sequencing batches were
     processed using the same DNA extraction methods and sequencing
     machines, however two different exome capture panels were used,
     as described below. Genotyping was performed using a SNP
     genotyping array for WES1 through WES4 and using
     &quot;genotyping-by-sequencing&quot; (GxS) for WES5 through WES9. The
     first four sequencing batches were sequenced at Regeneron using
     custom NEB/Kapa reagents with the IDT (Integrated DNA
     Technologies) xGen capture platform, including custom exome
     capture regions. Samples starting with batch WES5 were
     sequenced using the Twist Bioscience Human
     Comprehensive Exome panel, combined with spike-ins for
     sequencing genotyping sites (see Genotyping Methods), the full
     mitochondrial genome, and coverage boosted at selected sites
     for assaying clonal hematopoiesis of indeterminate potential
     (CHIP). SFARI performed SNV/indel calling via DeepVariant and
     GATK to generate gVCFs, pairwise relatedness inferred using
     PLINK v1.9 IBD estimates from common SNPs (AF &ge; 0.01, dbSNP
     v151) with &ge;15% relatedness flagged, and comprehensive
     individual- and family-level quality control executed using the
     internal GenomeCheckMate pipeline to exclude samples based on
     contamination (&ge;5%), insufficient coverage (&lt;20x in &lt;80% of
     targets), sex discordance, pedigree/IBD inconsistencies,
     unregistered relationships, unexpected duplicates, or excess
     relatedness, after which QC-passing individuals (selecting the
     most recent passing sample per person) were retained for
     variant calling and joint genotyping.
     </li>
 </ul>
 <p>
 The complete, runnable command history for downloading, counting and annotating the SFARI data,
 alongside every other cohort in this collection, is in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/varFreqs.txt" target="_blank">makeDoc file</a>;
 search it for &quot;SFARI SPARK&quot;. The scripts it calls, including
-<tt>sparkMergeVcfAddCounts.sh</tt> (allele counts) and <tt>mergeAndAnnotate.sh</tt> (merge plus
+<tt>sparkMergeVcfAddCounts.sh</tt> (allele counts), <tt>sparkWgs45kToVcf.sh</tt> (45k WGS frequency table conversion),
+<tt>sparkWgs45kPvcfToSites.sh</tt> (45k WGS genotype counts) and <tt>mergeAndAnnotate.sh</tt> (merge plus
 consequence annotation), are in the
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/varFreqs" target="_blank">varFreqs scripts directory</a>.
 </p>
 
 <h2>References</h2>
 <p>
 SPARK Consortium. Electronic address: <A HREF="mailto:&#112;f&#101;&#108;&#105;&#99;&#105;&#97;&#110;o&#64;&#115;&#105;&#109;&#111;&#110;&#115;f&#111;&#117;&#110;&#100;a&#116;&#105;&#111;&#110;.&#111;&#114;g">&#112;f&#101;&#108;&#105;&#99;&#105;&#97;&#110;o&#64;&#115;&#105;&#109;&#111;&#110;&#115;f&#111;&#117;&#110;&#100;a&#116;&#105;&#111;&#110;.&#111;&#114;g</A><!-- above address is pfeliciano at simonsfoundation.org -->, SPARK Consortium.
 <a href="https://linkinghub.elsevier.com/retrieve/pii/S0896-6273(18)30018-7" target="_blank">
 SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research</a>.
 <em>Neuron</em>. 2018 Feb 7;97(3):488-493.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/29420931" target="_blank">29420931</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7444276/" target="_blank">PMC7444276</a>
 </p>