1bab8758e7a5fe01120847c7fa68de71d1aa709e
max
  Wed Sep 23 16:16:44 2026 -0700
commenting out a docs piece before release

diff --git src/hg/makeDb/trackDb/human/hg38/problematic.html src/hg/makeDb/trackDb/human/hg38/problematic.html
index 3af24a16c4e..ac090b7f24f 100644
--- src/hg/makeDb/trackDb/human/hg38/problematic.html
+++ src/hg/makeDb/trackDb/human/hg38/problematic.html
@@ -1,220 +1,222 @@
 <h2>Description</h2>
 
 <p>
 This container track helps call out sections of the genome that often cause problems or
 confusion when working with the genome. The hg19 genome has a track with the same name, but with
 more subtracks, as the GeT-RM and Genome-in-a-Bottle artifact variants do not exist 
 for hg38.
 
 <h3>Problematic Regions</h3>
 <p>
 The <b>Problematic Regions</b> track contains the following subtracks:
 <ul>
 <li>
 The <b>UCSC Unusual Regions</b> subtrack contains annotations collected at UCSC, 
 put together from other tracks, our experiences and support email list
 requests over the years. For example, it contains the most well-known gene
 clusters (IGH, IGL, PAR1/2, TCRA, TCRB, etc) and annotations for the GRC
 <a href="/cgi-bin/hgTracks?db=hg19&chromInfoPage=">fixed sequences, alternate haplotypes, unplaced
 contigs, pseudo-autosomal regions, and mitochondria</a>. These loci can yield alignments with
 low-quality mapping scores and discordant read pairs, especially for short-read sequencing data.
 The data set was manually curated, based on the <a href="/cgi-bin/hgGateway">Genome Browser's
 assembly</a> description, the <a href="/FAQ/FAQdownloads.html">FAQs</a> about assembly, and the
 <a href="/cgi-bin/hgTrackUi?db=hg19&g=refSeqComposite">NCBI RefSeq &quot;other&quot; annotations</a>
 track data.
 </li>
 
 <li>
 The <b>ENCODE Blacklist</b> subtrack contains a comprehensive set of regions which are troublesome
 for high-throughput Next-Generation Sequencing (NGS) aligners. These regions tend to have a very
 high ratio of multi-mapping to unique mapping reads and high variance in mappability due to
 repetitive elements such as satellite, centromeric and telomeric repeats. 
 </li>
 
 <li>
 The <b>GRC Exclusions</b> subtrack contains a set of regions that have been flagged by the GRC to
 contain false duplications or contamination sequences. The GRC has now removed these sequences from
 the files that it uses to generate the reference assembly, however, removing the sequences from the
 GRCh38/hg38 assembly would trigger the next major release of the human assembly. In order to
 help users recognize these regions and avoid them in their analyses, the GRC have produced a masking
 file to be used as a companion to GRCh38, and the BED file is available from the
 <a href="https://ftp.ncbi.nlm.nih.gov/genomes/all/GCF/000/001/405/GCF_000001405.39_GRCh38.p13/GRCh38_major_release_seqs_for_alignment_pipelines/GCA_000001405.15_GRCh38_GRC_exclusions.bed"
 target="_blank">GenBank FTP site</a>.
 </li>
 </ul>
 
 <h3>Highly Reproducible Regions (HighRepro)</h3>
 <p>
 The <b>Highly Reproducible Regions</b> track highlights regions and variants
 from eight samples that can be used to assess variant detection pipelines. The
 &quot;Highly Reproducible Regions&quot; subtrack comprises the intersection of the reproducible
 regions across all eight samples, while the &quot;Variants&quot; subtracks contain the reproducible
 variants from each assayed sample. Both tracks contain data from the following samples:
 </p>
 <ul>
   <li>a Chinese Quartet, samples <b>CQ-5</b>, <b>CQ-6</b>, <b>CQ-7</b>, <b>CQ-8</b></li>
   <li>a HapMap Trio, samples <b>NA10385</b>, <b>NA12248</b>, <b>NA12249</b></li>
   <li>a Genome in a Bottle sample, <b>NA12878s</b></li>
 </ul>
 
 Please refer to the <em>Pan et al</em> reference for more information on how
 these regions were defined.
 </p>
 
 <h3>GIAB Problematic Regions</h3>
 <p>The <b>Genome in a Bottle (GIAB) Problematic Regions</b> tracks provide stratifications of the
 genome to evaluate variant calls in complex regions. It is designed for use with Global Alliance
 for Genomic Health (GA4GH) benchmarking tools like
 <a href="https://github.com/Illumina/hap.py" target="_blank">hap.py</a>
 and includes regions with low complexity, segmental duplications, functional regions,
 and difficult-to-sequence areas. Developed in collaboration with GA4GH, the
 <a href="https://www.nist.gov/programs-projects/genome-bottle"
 target="_blank">Genome in a Bottle (GIAB) consortium</a>, and the
 <a href="https://sites.google.com/ucsc.edu/t2tworkinggroup"
 target="_blank">Telomere-to-Telomere Consortium (T2T)</a>, the dataset aims to standardize the
 analysis of genetic variation by offering pre-defined BED files for stratifying true and false
 positives in genomic studies, facilitating accurate assessments in complex areas of the genome.</p>
 
 <p>
 The creation of the GIAB Problematic Regions tracks involves using a pipeline and configuration to
 generate stratification BED files that categorize genomic regions based on specific challenges,
 such as low complexity or difficult mapping, to facilitate accurate benchmarking of variant calls.
 For more information on the pipeline and configuration used, please visit the following webpage:
 <a href="https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/genome-stratifications/v3.5/README.md">
 https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/genome-stratifications/v3.5/README.md</a>.
 If you have questions or comments, please write to Justin Zook (jzook@nist.gov).</p>
 
 <h3>Panmask Easy 151b Regions</h3>
 <p>
 The <b>Panmask Easy 151b Regions</b> subtrack contains a set of sample-agnostic easy regions where
 short-read variant calling reaches high accuracy. Easy regions are derived for variant filtration
 agnostic to individual samples. They are genomic intervals where general variant callers achieve
 high accuracy without sophisticated filtering.</p>
 <p>
 A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that
 could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples
 were inspected to exclude regions with excessively high or low coverage or those enriched with
 low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure
 are referred to as pm151:lenient, where &quot;pm&quot; stands for panmask. In addition, low
 complexity regions identified by SDUST were removed.</p>
 <p>The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and
 other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38,
 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter
 variant calls for clinical or research human samples. Like the HighRepro track in this container
 (see above), it shows regions that are easy to sequence, not those that are problematic. The data
 was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.</p>
 
+<!--
 <h3>Panmask Difficult 151b Regions</h3>
 <p>
 The <b>Panmask Difficult 151b Regions</b> subtrack is the complement of the Panmask Easy 151b Regions
 track above: it marks the bases of the genome that Panmask did not classify as easy, so a
 variant caller can expect lower accuracy there. Unlike Panmask Easy, this track directly
 represents difficult regions, matching the rest of this container track, so it is the version
 used in the Problematic Regions Recommended Track Set. It was built at UCSC, not downloaded, by
 inverting the Panmask Easy 151b regions with the <tt>featureBits</tt> tool and removing assembly
 gaps, restricted to the 24 chromosomes that Panmask itself covers.</p>
+-->
 
 <h2>Display Conventions and Configuration</h2>
 
 <p>
 Each track contains a set of regions of varying length with no special configuration options. 
 The <em>UCSC Unusual Regions</em> track has a mouse-over description, all other tracks have at most
 a name field, which can be shown in pack mode. The tracks are usually kept in dense mode.
 </p>
 
 <p>
 The <em>Hide empty subtracks</em> control hides subtracks with no data in the browser window.
 Changing the browser window by zooming or scrolling may result in the display of a different
 selection of tracks.
 </p>
 
 <H2>Data access</H2>
 <p>
 The raw data can be explored interactively with the <a href="../cgi-bin/hgTables">Table Browser</a>
 or the <a href="../cgi-bin/hgIntegrator">Data Integrator</a>.
 
 <p>
 For automated download and analysis, the genome annotation is stored in bigBed files that
 can be downloaded from
 <a href="http://hgdownload.soe.ucsc.edu/gbdb/$db/bbi/problematic/" target="_blank">our download server</a>.
 Individual
 regions or the whole genome annotation can be obtained using our tool <tt>bigBedToBed</tt>
 which can be compiled from the source code or downloaded as a precompiled
 binary for your system. Instructions for downloading source code and binaries can be found
 <a href="http://hgdownload.soe.ucsc.edu/downloads.html#utilities_downloads">here</a>.
 The tool
 can also be used to obtain only features within a given range, e.g. 
 <br>
 <tt>bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/problematic/comments.bb -chrom=chr21 -start=0 -end=100000000 stdout</tt></p>
 </p>
 
 <p>
 <h2>Methods</h2>
 
 <p>
 Files were downloaded from the respective databases and converted to bigBed format.
 The procedure is documented in our
 <a href="https://github.com/ucscGenomeBrowser/kent/blob/master/src/hg/makeDb/doc/hg38/problematic.txt"
 target="_blank">hg38 makeDoc file</a>.
 </p>
 
 <h2>Credits</h2>
 <p>
 Thanks to Anna Benet-Pag&egrave;s, Max Haeussler, Angie Hinrichs, Daniel Schmelter, and Jairo
 Navarro at the UCSC Genome Browser for planning, building, and testing these tracks. The
 underlying data comes from the
 <a href="https://github.com/Boyle-Lab/Blacklist/blob/master/lists/hg19-blacklist-README.pdf"
 target="_blank">ENCODE Blacklist</a> and some parts were copied manually from the HGNC and NCBI
 RefSeq tracks.
 </p>
 
 <h2>References</h2>
 <p>
 Amemiya HM, Kundaje A, Boyle AP.
 <a href="https://www.nature.com/articles/s41598-019-45839-z" target="_blank">
 The ENCODE Blacklist: Identification of Problematic Regions of the Genome</a>.
 <em>Sci Rep</em>. 2019 Jun 27;9(1):9354.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/31249361" target="_blank">31249361</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6597582/" target="_blank">PMC6597582</a>
 </p>
 
 <p>
 Dwarshuis N, Kalra D, McDaniel J, Sanio P, Alvarez Jerez P, Jadhav B, Huang WE, Mondal R, Busby B,
 Olson ND <em>et al</em>.
 <a href="https://doi.org/10.1038/s41467-024-53260-y" target="_blank">
 The GIAB genomic stratifications resource for human reference genomes</a>.
 <em>Nat Commun</em>. 2024 Oct 19;15(1):9029.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/39424793" target="_blank">39424793</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11489684/" target="_blank">PMC11489684</a>
 </p>
 
 <p>
 Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, Gonzalez-Porta M, Eberle MA,
 Tezak Z, Lababidi S <em>et al</em>.
 <a href="https://doi.org/10.1038/s41587-019-0054-x" target="_blank">
 Best practices for benchmarking germline small-variant calls in human genomes</a>.
 <em>Nat Biotechnol</em>. 2019 May;37(5):555-560.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/30858580" target="_blank">30858580</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6699627/" target="_blank">PMC6699627</a>
 </p>
 
 <p>
 Li H.
 <a href="https://pmc.ncbi.nlm.nih.gov/articles/pmid/40799803/" target="_blank">
 Finding easy regions for short-read variant calling from pangenome data</a>.
 <em>ArXiv</em>. 2025 Aug 8;.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/40799803" target="_blank">40799803</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12340882/" target="_blank">PMC12340882</a>
 </p>
 
 <p>
 Pan B, Ren L, Onuchic V, Guan M, Kusko R, Bruinsma S, Trigg L, Scherer A, Ning B, Zhang C <em>et
 al</em>.
 <a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02569-8"
 target="_blank">
 Assessing reproducibility of inherited variants detected with short-read whole genome
 sequencing</a>.
 <em>Genome Biol</em>. 2022 Jan 3;23(1):2.
 PMID: <a href="https://www.ncbi.nlm.nih.gov/pubmed/34980216" target="_blank">34980216</a>; PMC: <a
 href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8722114/" target="_blank">PMC8722114</a>
 </p>