a5c699a7301156154700f51a20ae571bc6987051
max
  Sat Sep 26 17:53:41 2026 -0700
varFreqs: remove the AF-table-based sfariSparkWgs45k subtrack, superseded by the genotype-based sfariSparkWgs45kAsd; drop its scripts, the DSCAM demo script and their makeDoc sections, refs #38424

diff --git src/hg/makeDb/trackDb/human/sfariSparkExomes.html src/hg/makeDb/trackDb/human/sfariSparkExomes.html
index 93aae17eccc..d22015a1b5f 100644
--- src/hg/makeDb/trackDb/human/sfariSparkExomes.html
+++ src/hg/makeDb/trackDb/human/sfariSparkExomes.html
@@ -1,35 +1,33 @@
 <h2>Description</h2>
 <p>
 The <a href="https://sparkforautism.org/" target="_blank">Simons Foundation Autism Research
 Initiative (SFARI)</a> recruited a large cohort of families with autistic children who provided
 DNA samples and phenotypes. 54,558 families, parents and their children were sequenced, a total
 of 142,357 individuals with whole-exome (WES) and 12,519 with whole-genome sequencing (WGS).
 The data contains 32,559 trios and 8,895 quads (one sibling without autism), and 824 twins.
 </p>
 
 <p>
 The <a href="https://base.sfari.org/dataset/DS0000135" target="_blank">SPARK WGS August 2026
 release</a> (SFARI Base dataset DS0000135) adds whole genomes for 45,178 individuals from
 21,003 families, 20,858 of them with autism. This is a mostly new set of people: only 14 of
 them are also in the 12,519-genome release and 201 in the exome release. It includes about
 4,900 trios with an autistic child and 1,800 quads. The genomes were sequenced PCR-free on
-the Illumina NovaSeq X at Broad Clinical Labs. Two tracks show this release. "SFARI SPARK 45k
-WGS" was made from a frequency table that SFARI released with the data. It has no
-autism/non-autism split, its allele counts are estimates, and variants seen only once are
-left out. "SFARI SPARK 45k WGS ASD" was computed from the genotypes. It has exact counts,
-the autism/non-autism split and all variants, including those seen once (see Methods).
+the Illumina NovaSeq X at Broad Clinical Labs. The track "SFARI SPARK 45k WGS ASD" shows
+this release. Its counts were computed from the genotypes, with the autism/non-autism split
+and all variants, including those seen only once (see Methods).
 </p>
 
 <p>
 The same frequencies shown here are also available publicly on the
 <a href="https://genomes.sfari.org/" target="_blank">SFARI Genome Browser</a>.
 See (SPARK et al, Neuron 2018) for details.
 </p>
 
 <h3>Phenotype-stratified counts</h3>
 <p>
 In addition to the overall allele count (AC), allele number (AN), and allele
 frequency (AF), each variant record carries counts split by autism status
 (the <tt>asd</tt> column of the SPARK individual registration file):
 </p>
 <ul>
@@ -68,48 +66,32 @@
 downloaded from <a href="https://base.sfari.org/" target="_blank">SFARI Base</a> using Globus.
 No minimum allele frequency cutoff was applied.</p>
 
 <p>
 Because the genotype-level pVCFs cannot be redistributed, they were reduced to anonymous,
 sites-only VCFs carrying only the overall allele count (AC), allele number (AN) and frequency
 (AF), plus the autism-status counts described above, with the bcftools
 <a href="http://samtools.github.io/bcftools/howtos/plugin.fill-tags.html" target="_blank">fill-tags</a>
 plugin (its <tt>-S</tt> option produces the ASD/non-ASD splits), then normalized. The variants
 are also annotated with predicted protein consequences using
 <a href="http://samtools.github.io/bcftools/bcftools.html#csq" target="_blank">bcftools csq</a>
 against Ensembl gene models; that annotation is displayed in the combined frequency tracks of
 this collection. The exact commands for both steps are in the makeDoc file linked below.</p>
 
 <p>
-The WGS August 2026 release is handled differently. It was downloaded from SFARI Base with
-Globus, but only the table <tt>SPARK.WGS.2026_08.gatk.pvcf_variant_frequencies.tsv</tt> was
-used. SFARI made this table from the pVCFs with <tt>bcftools query</tt>. It lists chromosome,
-position, REF, ALT and AF, with AF rounded to six decimals. It has no allele count or allele
-number, so we estimated them. At every site, including sites on chrX and chrY, all low
-frequencies are exact multiples of 1/90,356, and 90,356 is two alleles for each of the
-45,178 individuals. For example, all of the 179 million alleles seen once have AF=1.1e-05.
-None has 1.2e-05, which is what a site with more than about 4% missing genotypes would give.
-We therefore set AN=90,356 for every variant and computed AC as AF &times; AN, rounded to the
-nearest integer. Multiallelic records were split into one record per alternate allele.
-Alleles with an estimated AC of 1 were removed. The remaining alleles were left-aligned
-against the reference with <tt>bcftools norm</tt>. The script is
-<tt>sparkWgs45kToVcf.sh</tt>.
-</p>
-
-<p>
-The track "SFARI SPARK 45k WGS ASD" was made from the genotype-level pVCFs of the same
-release, which are split into 2.5 Mb chunks. The samples were grouped by the <tt>asd</tt>
+The track "SFARI SPARK 45k WGS ASD" was made from the genotype-level pVCFs of the WGS August
+2026 release, downloaded from SFARI Base with Globus. They are split into 2.5 Mb chunks. The samples were grouped by the <tt>asd</tt>
 column of the release's sample metadata file: 20,858 autistic and 24,320 non-autistic
 individuals, with no sample left unassigned. As for the older SPARK tracks, <tt>bcftools
 +fill-tags</tt> computed AC, AN and AF overall and per group from the called genotypes.
 Missing genotypes do not count towards AN. The genotypes were then removed. Multiallelic
 records were split and left-aligned with <tt>bcftools norm</tt>. Alleles that no individual
 carries after joint genotyping (AC=0) were removed. There is no allele count cutoff.
 Records that GLnexus marks with the filter <tt>MONOALLELIC</tt> are alleles it could not
 merge into an overlapping site. In these, only the carriers have a genotype, so AN would
 count only them and AF would be 1. For these records, AN was set to twice the number of
 individuals in each group and AF was recomputed from AC. The INFO field <tt>VARLEN</tt> is
 the length of ALT minus the length of REF: positive for insertions, negative for
 deletions, 0 for substitutions. The pVCFs were processed in 500 kb pieces on our compute
 cluster. The scripts are <tt>sparkWgs45kPvcfToSites.sh</tt>,
 <tt>sparkWgs45kPvcfJobs.sh</tt> and <tt>sparkWgs45kPvcfMerge.sh</tt>.
 </p>
@@ -182,30 +164,30 @@
     individual- and family-level quality control executed using the
     internal GenomeCheckMate pipeline to exclude samples based on
     contamination (&ge;5%), insufficient coverage (&lt;20x in &lt;80% of
     targets), sex discordance, pedigree/IBD inconsistencies,
     unregistered relationships, unexpected duplicates, or excess
     relatedness, after which QC-passing individuals (selecting the
     most recent passing sample per person) were retained for
     variant calling and joint genotyping.
     </li>
 </ul>
 <p>
 The complete, runnable command history for downloading, counting and annotating the SFARI data,
 alongside every other cohort in this collection, is in the
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/varFreqs.txt" target="_blank">makeDoc file</a>;
 search it for &quot;SFARI SPARK&quot;. The scripts it calls, including
-<tt>sparkMergeVcfAddCounts.sh</tt> (allele counts), <tt>sparkWgs45kToVcf.sh</tt> (45k WGS frequency table conversion),
+<tt>sparkMergeVcfAddCounts.sh</tt> (allele counts),
 <tt>sparkWgs45kPvcfToSites.sh</tt> (45k WGS genotype counts) and <tt>mergeAndAnnotate.sh</tt> (merge plus
 consequence annotation), are in the
 <a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/varFreqs" target="_blank">varFreqs scripts directory</a>.
 </p>
 
 <h2>References</h2>
 <p>
 SPARK Consortium. Electronic address: <A HREF="mailto:&#112;f&#101;&#108;&#105;&#99;&#105;&#97;&#110;o&#64;&#115;&#105;&#109;&#111;&#110;&#115;f&#111;&#117;&#110;&#100;a&#116;&#105;&#111;&#110;.&#111;&#114;g">&#112;f&#101;&#108;&#105;&#99;&#105;&#97;&#110;o&#64;&#115;&#105;&#109;&#111;&#110;&#115;f&#111;&#117;&#110;&#100;a&#116;&#105;&#111;&#110;.&#111;&#114;g</A><!-- above address is pfeliciano at simonsfoundation.org -->, SPARK Consortium.
 <a href="https://linkinghub.elsevier.com/retrieve/pii/S0896-6273(18)30018-7" target="_blank">
 SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research</a>.
 <em>Neuron</em>. 2018 Feb 7;97(3):488-493.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/29420931" target="_blank">29420931</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7444276/" target="_blank">PMC7444276</a>
 </p>