9cedfa38c14068c79dec89f76c606ee22b3f931a
max
  Wed Sep 30 15:02:37 2026 -0700
phasedVars: new subtrack hgdp1kSnv, a 17GB version of the 3.5TB gnomAD HGDP+1000G genotype VCF with only SNVs with AC>5 and only GT, so haplotype clustering can be shown up to 5Mbp, refs #37306

diff --git src/hg/makeDb/trackDb/human/phasedVars.html src/hg/makeDb/trackDb/human/phasedVars.html
index 74fa8148711..80f7a4496df 100644
--- src/hg/makeDb/trackDb/human/phasedVars.html
+++ src/hg/makeDb/trackDb/human/phasedVars.html
@@ -95,30 +95,42 @@
 <p>
 <b>MXB:</b> Allele frequencies by geographical state and ancestry are available via
 the <a target="_blank" href="https://morenolab.shinyapps.io/mexvar/">MexVar platform</a>.
 Raw genotype data are available under controlled access at the
 EGA (Study: EGAS00001005797; Dataset: EGAD00010002361). For the VCFs, email
 andres.moreno@cinvestav.mx.
 </p>
 
 <h2>Methods</h2>
 <p>
 <b>SGDP:</b> The version used was
 <a target="_blank" href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/"
 >https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/</a>,
 merged with bcftools and lifted to hg38 with CrossMap.
 </p>
+<p>
+<b>gnomAD HGDP+1000G, SNVs AC&gt;5:</b> The full gnomAD callset is 3.5 TB and too slow
+to display in haplotype clustering mode except in very small windows. For this
+second, smaller version we kept only single-nucleotide variants with an allele count
+(INFO/AC) higher than five, only the GT genotype field and only the AC, AN and AF
+INFO fields, using bcftools. Indels, rare variants and all other fields are only in
+the full track. See the
+<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/varFreqs.txt"
+target="_blank">makeDoc</a> and the
+<a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/scripts/varFreqs/hgdp1kCommonSnvs.sh"
+target="_blank">script</a> for details.
+</p>
 
 <h2>Credits</h2>
 <p>
 <b>MXB:</b> We thank the Center for Research and Advanced Studies (Cinvestav) of Mexico for
 generating and providing the frequency data, the National Institute of Medical
 Sciences and Nutrition (INCMNSZ) for DNA extraction, and the Ministry of Health
 together with the National Institute of Public Health (INSP) for the design and
 implementation of the National Health Survey 2000 (ENSA 2000). We also thank
 the ENSA-Genomics Consortium for their contributions to sample collection and
 data processing that made possible the construction of the MXB genomic
 resource.
 </p>
 <p>
 <b>SGDP:</b> This project was funded by the Simons Foundation. Thanks to David Reich and Swapan 
 Mallick for help with importing the data.