8d5759442b9034d07823c7c57ad4c53f7fc2d6bc
max
  Sat Jul 25 18:25:07 2026 -0700
sfariSparkExomes: update track description

diff --git src/hg/makeDb/trackDb/human/sfariSparkExomes.html src/hg/makeDb/trackDb/human/sfariSparkExomes.html
index 5a710f64574..1040d98c369 100644
--- src/hg/makeDb/trackDb/human/sfariSparkExomes.html
+++ src/hg/makeDb/trackDb/human/sfariSparkExomes.html
@@ -39,40 +39,46 @@
 </p>
 <p>
 Allele frequencies can also be displayed on the
 <a href="https://genomes.sfari.org/" target="_blank">SFARI Genome Browser</a>.
 Full CRAMs and VCFs with genotypes are available from
 <a href="https://base.sfari.org/" target="_blank">SFARI Base</a>.
 They require a data access request, which is usually reviewed quickly. More information is
 available in the
 <a href="https://cohorts-cdn.simonsfoundation.org/spark/researcher_packets/SPARK_SFARI_Researcher_Welcome_Packet.pdf"
 target="_blank">SPARK Welcome Packet</a>.
 </p>
 
 <h2>Methods</h2>
 
 <p>The genome browser track project was approved by the Simons Foundation under request
-number 14584.1. WES and WGS data were downloaded from
-<a href="https://base.sfari.org/" target="_blank">SFARI Base</a>.
-pVCFs were downloaded, anonymized with a script using bcftools and its &quot;fill-tags&quot; plugin and
-normalized. There was no minimum allele frequency cutoff.
-The ASD-status sample-group file derived from the SPARK <tt>individuals_registration</tt>
-TSV was passed to fill-tags via its <tt>-S</tt> option, which adds the per-group
-<tt>AC_AUT</tt>/<tt>AN_AUT</tt>/<tt>AF_AUT</tt> and <tt>AC_NON_AUT</tt>/<tt>AN_NON_AUT</tt>/<tt>AF_NON_AUT</tt>
-tags alongside the overall AC/AN/AF.</p>
+number 14584.1. The multi-sample project VCFs (pVCFs) for both the WES and WGS releases were
+downloaded from <a href="https://base.sfari.org/" target="_blank">SFARI Base</a> using Globus.
+No minimum allele frequency cutoff was applied.</p>
+
+<p>
+Because the genotype-level pVCFs cannot be redistributed, they were reduced to anonymous,
+sites-only VCFs carrying only the overall allele count (AC), allele number (AN) and frequency
+(AF), plus the autism-status counts described above, with the bcftools
+<a href="http://samtools.github.io/bcftools/howtos/plugin.fill-tags.html" target="_blank">fill-tags</a>
+plugin (its <tt>-S</tt> option produces the ASD/non-ASD splits), then normalized. The variants
+are also annotated with predicted protein consequences using
+<a href="http://samtools.github.io/bcftools/bcftools.html#csq" target="_blank">bcftools csq</a>
+against Ensembl gene models; that annotation is displayed in the combined frequency tracks of
+this collection. The exact commands for both steps are in the makeDoc file linked below.</p>
 
-<p>The methods are documented as follows by SFARI:</p>
+<p>The sequencing and variant-calling methods are documented as follows by SFARI:</p>
 <ul>
   <li>
     <b>WGS</b>:
     This release consists of sequence and variant call data for 12,519
     unique individuals, of which 12,517 (99.98%) have available genome-wide
     SNP genotype data. Sequencing and genotyping of all samples in this
     release was performed at New York Genome Center (NYGC). DNA from saliva
     samples were extracted and prepared with PCR-free methods and sequenced
     with paired-end sequencing of 150 bases on the Illumina NovaSeq 6000
     system. Alignment of reads to the human reference genome version
     GRCh38, duplicate read marking, and Base Quality Score Recalibration
     (BQSR) were performed by New York Genome Center (NYGC). Whole-genome
     sequencing data were processed using a standardized, functionally
     equivalent CCDG pipeline with alignment to the GRCh38DH (1000 Genomes)
     reference using BWA-MEM v0.7.15 (deterministic settings, no -M, use of
@@ -115,28 +121,33 @@
     (CHIP). SFARI performed SNV/indel calling via DeepVariant and
     GATK to generate gVCFs, pairwise relatedness inferred using
     PLINK v1.9 IBD estimates from common SNPs (AF &ge; 0.01, dbSNP
     v151) with &ge;15% relatedness flagged, and comprehensive
     individual- and family-level quality control executed using the
     internal GenomeCheckMate pipeline to exclude samples based on
     contamination (&ge;5%), insufficient coverage (&lt;20x in &lt;80% of
     targets), sex discordance, pedigree/IBD inconsistencies,
     unregistered relationships, unexpected duplicates, or excess
     relatedness, after which QC-passing individuals (selecting the
     most recent passing sample per person) were retained for
     variant calling and joint genotyping.
     </li>
 </ul>
 <p>
-The <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/varFreqs.txt" target="_blank">makeDoc file</a> documents how all source files of the varFreqs track were converted.
-For some tracks, python scripts were necessary and are also available from <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/scripts/varFreqs" target="_blank">GitHub</a>.
+The complete, runnable command history for downloading, counting and annotating the SFARI data,
+alongside every other cohort in this collection, is in the
+<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/varFreqs.txt" target="_blank">makeDoc file</a>;
+search it for &quot;SFARI SPARK&quot;. The scripts it calls, including
+<tt>sparkMergeVcfAddCounts.sh</tt> (allele counts) and <tt>mergeAndAnnotate.sh</tt> (merge plus
+consequence annotation), are in the
+<a href="https://github.com/ucscGenomeBrowser/kent/tree/master/src/hg/makeDb/scripts/varFreqs" target="_blank">varFreqs scripts directory</a>.
 </p>
 
 <h2>References</h2>
 <p>
 SPARK Consortium. Electronic address: <A HREF="mailto:&#112;f&#101;&#108;&#105;&#99;&#105;&#97;&#110;o&#64;&#115;&#105;&#109;&#111;&#110;&#115;f&#111;&#117;&#110;&#100;a&#116;&#105;&#111;&#110;.&#111;&#114;g">&#112;f&#101;&#108;&#105;&#99;&#105;&#97;&#110;o&#64;&#115;&#105;&#109;&#111;&#110;&#115;f&#111;&#117;&#110;&#100;a&#116;&#105;&#111;&#110;.&#111;&#114;g</A><!-- above address is pfeliciano at simonsfoundation.org -->, SPARK Consortium.
 <a href="https://linkinghub.elsevier.com/retrieve/pii/S0896-6273(18)30018-7" target="_blank">
 SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research</a>.
 <em>Neuron</em>. 2018 Feb 7;97(3):488-493.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/29420931" target="_blank">29420931</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7444276/" target="_blank">PMC7444276</a>
 </p>