5fd99d465a31d5bd38320ac85ff1969ab9515014
mspeir
Wed Sep 23 15:44:06 2026 -0700
Regulation FAQ: correct assembly coverage, add variant effect section, refs #24610
Corrections found by checking the page against our own release announcements,
which I should have done before the first commit:
- ENCODE4 Regulation is on mm10 as well as hg38. The track symbol differs by
assembly (wgEncodeReg4 on hg38, encode4Reg on mm10), so a search by name
missed it.
- JASPAR 2026 covers hg38, mm39, danRer11, galGal6, dm6, ce11, ci3 and
sacCer3. hg19 and mm10 stop at JASPAR 2024. The page had implied all ten
were 2026. Also notes that the track holds several releases as subtracks,
so the version depends on which one is turned on.
- TFBS Conserved is on hg17 as well as hg18 and hg19.
- ENCODE4 replaced ENCODE3 as the default in July 2026 and ENCODE3 is kept
for archival use. Said in the text and marked in the summary table.
- The cell type section described the old subtrack list for the large
collections, which now use the faceted interface.
- Dropped "composite" and "superTrack", which users should not see.
New section on what a single variant does to regulation, covering MPRAVarDB
for variants that were actually tested in a reporter assay, and AlphaGenome
and PromoterAI for predictions. These live under Phenotype and Disease
Associations rather than Regulation, which is worth saying since that is
where a reader would otherwise fail to find them.
Links to goldenPath/help/hgRegMotifHelp.html from the motif definition and
from the JASPAR entry, and the definition now matches the wording there.
Co-Authored-By: Claude Opus 5 (1M context)
Return to FAQ Table of Contents
The assembly names after each track are links. They open that track's description page on
that assembly, which gives the methods, the data version and the citation. Coverage varies a
lot between assemblies, so check the list before you assume a track exists on the genome you
work with.Topics
The basics
@@ -41,32 +43,34 @@
them are in the Regulation track group, which you will find below the browser image on
the main browser page. It is not always obvious which
one you want, because these tracks answer quite different questions even when they all look
like boxes on the screen.
Nearly every question we get about these tracks comes down to this distinction.
A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.
-A predicted binding site comes from scanning the genome sequence for a motif, a short -pattern that the factor is known to prefer. Predictions exist everywhere in the genome for every +A predicted binding site comes from scanning the genome sequence for a +motif, the sequence theme a given factor +prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist +everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.
Neither kind is better than the other. To find out where a factor was actually found, use a measured track. To find out whether some sequence you care about, a variant or a promoter fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a prediction as evidence that the factor binds there.
For human, three tracks cover most needs. All three are in the Regulation group.
@@ -78,47 +82,51 @@ single project. There are three versions: all peaks per experiment, a non-redundant set that merges similar targets, and cis-regulatory modules. Available for hg19, hg38, mm10, mm39 and dm6.Two older tracks still come up in questions. TFBS Conserved shows predicted sites that are conserved across human, mouse and rat, on hg17, hg18 and hg19 only; it has not been updated in many years and there is no hg38 version. ORegAnno is a curated collection of regulatory elements taken from the literature, so it is small but every entry has a citation, on hg19, hg38, mm10 and dm6.
@@ -238,71 +246,87 @@ hg38.-For the chromatin evidence behind the called elements, see -ENCODE4 Regulation on hg38, which -carries DNase, ATAC-seq, histone -modification and CTCF signal organized by tissue, and the older ENCODE3 Regulation container on +For the chromatin evidence behind the called elements, see ENCODE4 Regulation, +which carries DNase, ATAC-seq, histone modification and CTCF signal organized by tissue, on +hg38 and +mm10. This replaced ENCODE3 +Regulation +as the default in July 2026; ENCODE3 is kept for archival use and is still on hg19 and hg38.
A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only and so on. The word candidate is doing real work here: these are regions worth testing, not confirmed regulatory elements.
-The ENCODE cCREs container on hg38 holds -several versions. The ENCODE4 -cCREs registry is the current one and should be your default. The ENCODE4 Core -Collection is a smaller, higher-confidence subset. ENCODE3 cCREs is the -earlier release, kept because a great many published analyses used it and coordinates need to -stay reproducible. The container also carries per-biosample subtracks, which is how you restrict -the classification to one cell type. Mouse -mm10 carries the same three versions.
+The ENCODE cCREs container on +hg38 holds several versions. The +ENCODE4 +cCREs registry became the default in July 2026 and is the one to use: 2.3 million human +and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, +higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with +all four core assays. ENCODE3 cCREs is the earlier release, kept for archival +use because a great many published analyses used it and coordinates need to stay reproducible. The +container also carries per-biosample subtracks, which is how you restrict +the classification to one cell type. Mouse +mm10 carries the same three versions.-Most of the large regulatory tracks are composites or superTracks holding many subtracks, one -per cell type or experiment, and they arrive with only a summary view turned on. Click the track -name to open its configuration page, where you will find the list of subtracks and, on the -bigger tracks, filters. ReMap, for instance, lets you filter by transcription factor directly in -the track settings, so you can show one factor across all its experiments.
+Most of the large regulatory tracks are collections holding many separate tracks, one per cell +type or experiment, and they arrive with only a summary view turned on. Click the track name to +open its configuration page, where you will find the list and, on the bigger collections, +filters. ReMap, for instance, lets you filter by transcription factor directly in the track +settings, so you can show one factor across all its experiments. ++The largest collections use a different configuration page. Where there are hundreds or +thousands of individual experiments, as in the per-experiment tracks of +ENCODE4 Regulation on +hg38, the page shows facets down the +left side and a +paginated table on the right: tick the tissue, assay or biosample you want and the table narrows +to the matching experiments. This is usually a faster way to reach one cell type than reading a +long list.
To find a track rather than configure one, the Track Search page searches track names and descriptions across the whole assembly, which is usually faster than reading through the track groups.
No single track answers this, so you have to work outward from the gene.
To do this systematically, the Data Integrator will intersect two or more tracks and return a table, and the Table Browser will do the same for a region or for a list of genes.
+ ++This is a different question from the rest of this page. Everything above annotates regions and +tells you what is where; the tracks here take one base change and tell you what it does. The +same measured and predicted distinction applies, so start with the one track that is +measured.
++MPRAVarDB holds 239,028 variants that were put +through a reporter assay and scored for allelic effect, drawn from 18 MPRA studies covering more +than 30 cell lines and more than 30 diseases or traits. The variants come from GWAS and eQTL +fine-mapping, from saturation mutagenesis of 20 disease-associated regulatory elements, and from +smaller focused screens, so the coverage is concentrated on loci people have already had reason +to care about. Items are colored by significance, dark red for FDR below 0.05. If your variant +is in here, you have an experimental answer and not a guess. It lives in the +MPRAs collection, alongside MPRA Base, which tests whole elements. +Available for hg38.
++Most variants will not be in it, since 239,028 tested positions is a very small part of the +genome. Two prediction tracks fill the gap by scoring every possible substitution, and neither +is in the Regulation group: look under Phenotype and Disease Associations, in +the Deleteriousness Predictions collection on +hg38.
++A high score from either says a model thinks the change is disruptive, not that anyone has +measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement +is the better evidence.
+-Considerably less, and this is a matter of what data exists rather than what we have loaded. The -large regulatory resources were built for human first and mouse second. JASPAR predictions are -the most widely available, since they only require the genome sequence and a motif, and are -present for zebrafish, fly, worm, chicken, sea squirt and yeast in addition to human and mouse. -ReMap and ORegAnno cover fly. Most of the remaining tracks described on this page are human and -mouse only.
+Much less. The limit is usually the data itself: these resources were built for human first, and +most never went further than mouse. JASPAR is the main exception, since it needs only the genome +sequence and a motif, so it covers zebrafish, fly, worm, chicken, sea squirt and yeast as well. +ReMap and ORegAnno also cover fly. Everything else described on this page is human and mouse +only.-Coverage also differs between assemblies of the same organism. On mouse, mm10 carries most of -the regulatory tracks while mm39 has only JASPAR, ReMap and VISTA, because several of the source -projects have not released mm39 versions. If a track you need is on mm10 but not mm39, the -LiftOver tool can convert coordinates between the two, -though you should check the result. Human has the same problem on a smaller scale: a few of -these tracks are still hg19 only, and TFBS Conserved was never rebuilt for hg38.
+Mouse is split awkwardly between its own assemblies. mm10 carries most of these tracks and mm39 +has only JASPAR, ReMap and VISTA, because several of the source projects never released mm39 +versions. If you need something that is on mm10 but not mm39, +LiftOver will convert the coordinates, though check the +result before trusting it. Human has a milder version of the same problem: a few tracks are +still hg19 only, and TFBS Conserved was never rebuilt for hg38.For assemblies not hosted at UCSC, or for tracks we do not carry, check the public hubs list, where other groups publish data through our browser.
Each assembly name below links to that track's description page on that assembly. A few tracks appear on additional genomes not listed here; use Track Search to check a genome that is not shown.
| hg19, hg38, mm10, mm39, dm6 | |||||
| TF ChIP (ENCODE 3 TFBS on hg19) | ENCODE 3 TF ChIP-seq peaks, around 340 factors in about 130 cell types | Measured | hg19, hg38 | ||
| JASPAR Transcription Factors | -Motif matches from the JASPAR CORE collection | +Motif matches from the JASPAR CORE collection; several releases as subtracks | Predicted | -hg19, - hg38, - mm10, + | JASPAR 2026 on hg38, mm39, danRer11, + galGal6, dm6, ce11, - galGal6, ci3, - sacCer3 | + sacCer3; up to JASPAR 2024 on + hg19 and + mm10
| TFBS Conserved | Conserved motif matches; not updated recently, no hg38 version | Predicted | hg17, hg18, hg19 | ||
| ORegAnno | Regulatory elements curated from the literature | Measured, curated | hg19, hg38, @@ -444,78 +511,102 @@ | Mixed, with predicted targets | hg19, hg38 |
| VISTA Enhancers | Elements tested individually in transgenic mouse assays | Measured, validated | hg19, hg38, mm10, mm39 | ||
| ENCODE cCREs | -Candidate elements classified from chromatin signal | +Candidate elements classified from chromatin signal; ENCODE4 is the default | Predicted from measured signal | hg38, mm10 | |
| RefSeq Functional Elements | NCBI curated non-coding functional elements | Measured, curated | hg38, mm10 | ||
| MPRAs | -Massively parallel reporter assay activity | +MPRAs (MPRA Base) | +Reporter assay activity for 40,938 tested regulatory elements | Measured | hg38 |
| Chromatin and expression context | |||||
| ENCODE4 Regulation | -DNase, ATAC, histone marks and CTCF by tissue | +DNase, ATAC, histone marks and CTCF by tissue; the current default | Measured | -hg38 | +hg38, + mm10 |
| ENCODE3 Regulation | -DNase, histone marks and transcription signal | +DNase, histone marks and transcription signal; archival, replaced by ENCODE4 | Measured | hg19, hg38 | |
| Single-cell ATAC-seq | Accessibility peaks and signal from Cell Browser datasets | Measured | hg38, mm10 | ||
| GTEx Gene | Gene expression across 53 tissues | Measured | hg19, hg38 | ||
| GTEx cis-eQTLs | Variants associated with expression of nearby genes | Measured | hg38 | ||
| Effect of a single variant | +|||||
| MPRAVarDB | +239,028 variants tested for allelic effect in reporter assays | +Measured | +hg38 | +||
| AlphaGenome | +Variant Impact score for every single-base substitution, coding and non-coding; + under Phenotype and Disease Associations | +Predicted | +hg38 | +||
| PromoterAI | +Score for every single-base substitution in proximal promoters; under Phenotype and + Disease Associations | +Predicted | +hg38 | +||
For the full set of tracks on any assembly, open the Regulation group on the browser page, or use Track Search.