b197e1670a076b65838fd91b869a4ec1c3096c0f max Wed Sep 23 16:08:08 2026 -0700 Add Panmask Difficult 151b, the inverse of Panmask Easy 151b, for the Problematic Regions RTS Panmask marks easy regions under a "Problematic Regions" container, which Anna flagged as confusing. Rather than change the released Panmask Easy track, add a second track with the complement regions, built with featureBits (excluding assembly gaps and restricted to the 24 chromosomes Panmask itself covers). Checked the source first: Zenodo record 16755940 is still v1.4, same version already in use, MD5 verified. New track is alpha only for QA to pick up. refs #38375 diff --git src/hg/makeDb/trackDb/human/hg38/problematic.html src/hg/makeDb/trackDb/human/hg38/problematic.html index 20b2f96dbd3..3af24a16c4e 100644 --- src/hg/makeDb/trackDb/human/hg38/problematic.html +++ src/hg/makeDb/trackDb/human/hg38/problematic.html @@ -92,30 +92,40 @@ high accuracy without sophisticated filtering.</p> <p> A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples were inspected to exclude regions with excessively high or low coverage or those enriched with low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure are referred to as pm151:lenient, where "pm" stands for panmask. In addition, low complexity regions identified by SDUST were removed.</p> <p>The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38, 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter variant calls for clinical or research human samples. Like the HighRepro track in this container (see above), it shows regions that are easy to sequence, not those that are problematic. The data was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.</p> +<h3>Panmask Difficult 151b Regions</h3> +<p> +The <b>Panmask Difficult 151b Regions</b> subtrack is the complement of the Panmask Easy 151b Regions +track above: it marks the bases of the genome that Panmask did not classify as easy, so a +variant caller can expect lower accuracy there. Unlike Panmask Easy, this track directly +represents difficult regions, matching the rest of this container track, so it is the version +used in the Problematic Regions Recommended Track Set. It was built at UCSC, not downloaded, by +inverting the Panmask Easy 151b regions with the <tt>featureBits</tt> tool and removing assembly +gaps, restricted to the 24 chromosomes that Panmask itself covers.</p> + <h2>Display Conventions and Configuration</h2> <p> Each track contains a set of regions of varying length with no special configuration options. The <em>UCSC Unusual Regions</em> track has a mouse-over description, all other tracks have at most a name field, which can be shown in pack mode. The tracks are usually kept in dense mode. </p> <p> The <em>Hide empty subtracks</em> control hides subtracks with no data in the browser window. Changing the browser window by zooming or scrolling may result in the display of a different selection of tracks. </p> <H2>Data access</H2>