f03f56cd3c795a6fba2b8419662a9a2c5d49f69a
max
  Sat Sep 26 14:16:17 2026 -0700
sfariSparkWgs45kAsd: now genome-wide (518M variants from the 45,178 genotype pVCFs, run on parasol); per-allele INFO fields declared Number=1 so the VCF track filters accept them, doc page no longer says DSCAM only, refs #38424

diff --git src/hg/makeDb/trackDb/human/sfariSparkExomes.html src/hg/makeDb/trackDb/human/sfariSparkExomes.html
index b833859acbe..93aae17eccc 100644
--- src/hg/makeDb/trackDb/human/sfariSparkExomes.html
+++ src/hg/makeDb/trackDb/human/sfariSparkExomes.html
@@ -5,32 +5,31 @@
 DNA samples and phenotypes. 54,558 families, parents and their children were sequenced, a total
 of 142,357 individuals with whole-exome (WES) and 12,519 with whole-genome sequencing (WGS).
 The data contains 32,559 trios and 8,895 quads (one sibling without autism), and 824 twins.
 </p>
 
 <p>
 The <a href="https://base.sfari.org/dataset/DS0000135" target="_blank">SPARK WGS August 2026
 release</a> (SFARI Base dataset DS0000135) adds whole genomes for 45,178 individuals from
 21,003 families, 20,858 of them with autism. This is a mostly new set of people: only 14 of
 them are also in the 12,519-genome release and 201 in the exome release. It includes about
 4,900 trios with an autistic child and 1,800 quads. The genomes were sequenced PCR-free on
 the Illumina NovaSeq X at Broad Clinical Labs. Two tracks show this release. "SFARI SPARK 45k
 WGS" was made from a frequency table that SFARI released with the data. It has no
 autism/non-autism split, its allele counts are estimates, and variants seen only once are
 left out. "SFARI SPARK 45k WGS ASD" was computed from the genotypes. It has exact counts,
-the autism/non-autism split and all variants, including those seen once. For now it covers
-only the DSCAM gene on chr21 (see Methods).
+the autism/non-autism split and all variants, including those seen once (see Methods).
 </p>
 
 <p>
 The same frequencies shown here are also available publicly on the
 <a href="https://genomes.sfari.org/" target="_blank">SFARI Genome Browser</a>.
 See (SPARK et al, Neuron 2018) for details.
 </p>
 
 <h3>Phenotype-stratified counts</h3>
 <p>
 In addition to the overall allele count (AC), allele number (AN), and allele
 frequency (AF), each variant record carries counts split by autism status
 (the <tt>asd</tt> column of the SPARK individual registration file):
 </p>
 <ul>
@@ -92,36 +91,39 @@
 We therefore set AN=90,356 for every variant and computed AC as AF &times; AN, rounded to the
 nearest integer. Multiallelic records were split into one record per alternate allele.
 Alleles with an estimated AC of 1 were removed. The remaining alleles were left-aligned
 against the reference with <tt>bcftools norm</tt>. The script is
 <tt>sparkWgs45kToVcf.sh</tt>.
 </p>
 
 <p>
 The track "SFARI SPARK 45k WGS ASD" was made from the genotype-level pVCFs of the same
 release, which are split into 2.5 Mb chunks. The samples were grouped by the <tt>asd</tt>
 column of the release's sample metadata file: 20,858 autistic and 24,320 non-autistic
 individuals, with no sample left unassigned. As for the older SPARK tracks, <tt>bcftools
 +fill-tags</tt> computed AC, AN and AF overall and per group from the called genotypes.
 Missing genotypes do not count towards AN. The genotypes were then removed. Multiallelic
 records were split and left-aligned with <tt>bcftools norm</tt>. Alleles that no individual
-carries after joint genotyping (AC=0) were removed. There is no allele count cutoff. The
-INFO field <tt>VARLEN</tt> is the length of ALT minus the length of REF: positive for
-insertions, negative for deletions, 0 for substitutions. Unlike the other SPARK tracks,
-this track also shows the predicted protein consequences (<tt>BCSQ</tt>) on its details
-page, computed with <tt>bcftools csq</tt> against Ensembl release 115. The scripts are
-<tt>sparkWgs45kPvcfToSites.sh</tt> and <tt>sparkWgs45kPvcfRange.sh</tt>.
+carries after joint genotyping (AC=0) were removed. There is no allele count cutoff.
+Records that GLnexus marks with the filter <tt>MONOALLELIC</tt> are alleles it could not
+merge into an overlapping site. In these, only the carriers have a genotype, so AN would
+count only them and AF would be 1. For these records, AN was set to twice the number of
+individuals in each group and AF was recomputed from AC. The INFO field <tt>VARLEN</tt> is
+the length of ALT minus the length of REF: positive for insertions, negative for
+deletions, 0 for substitutions. The pVCFs were processed in 500 kb pieces on our compute
+cluster. The scripts are <tt>sparkWgs45kPvcfToSites.sh</tt>,
+<tt>sparkWgs45kPvcfJobs.sh</tt> and <tt>sparkWgs45kPvcfMerge.sh</tt>.
 </p>
 
 <p>The sequencing and variant-calling methods are documented as follows by SFARI:</p>
 <ul>
   <li>
     <b>WGS, August 2026 release</b>: DNA was extracted from saliva at Broad Clinical Labs or
     Prevention Genetics. Genomes were sequenced PCR-free on the Illumina NovaSeq X at Broad
     Clinical Labs. Samples were kept if contamination was 2.5% or less and 75% or more of
     reads were mapped. Genetic relationships were checked against the SPARK pedigrees. Reads
     were realigned to GRCh38 with BWA-MEM2 v2.3, duplicates were marked with GATK
     MarkDuplicates, and per-sample gVCFs were called with GATK HaplotypeCaller v4.6.2.0. All
     gVCFs were then joint-called with GLnexus v1.4.1 using its GATK configuration.
   </li>
   <li>
     <b>WGS, 12,519 genomes</b>: