2ae817bcd648f2a3184bb4b118840e65749fea2d max Thu Sep 24 04:15:20 2026 -0700 problematic.txt/html: reconcile Panmask Easy/Difficult coverage to 87.8%/12.2%, one decimal place, refs #38375 diff --git src/hg/makeDb/trackDb/human/hg38/problematic.html src/hg/makeDb/trackDb/human/hg38/problematic.html index ac090b7f24f..66834aac902 100644 --- src/hg/makeDb/trackDb/human/hg38/problematic.html +++ src/hg/makeDb/trackDb/human/hg38/problematic.html @@ -86,31 +86,31 @@

Panmask Easy 151b Regions

The Panmask Easy 151b Regions subtrack contains a set of sample-agnostic easy regions where short-read variant calling reaches high accuracy. Easy regions are derived for variant filtration agnostic to individual samples. They are genomic intervals where general variant callers achieve high accuracy without sophisticated filtering.

A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples were inspected to exclude regions with excessively high or low coverage or those enriched with low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure are referred to as pm151:lenient, where "pm" stands for panmask. In addition, low complexity regions identified by SDUST were removed.

The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and -other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38, +other genomic regions where short-read mapping is often problematic. They cover 87.8% of hg38, 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter variant calls for clinical or research human samples. Like the HighRepro track in this container (see above), it shows regions that are easy to sequence, not those that are problematic. The data was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.