b3aef6de6abdf2356bf7d4a64fd5d857726eea6b
mspeir
Sun Oct 4 09:57:49 2026 -0700
Regulation FAQ: cis-regulation background and six interpretive questions
Max asked for a short introduction to cis-regulation, and Lou noted that the
page leaned towards listing datasets rather than explaining how to read them.
This covers both.
A new opening question walks through the kinds of evidence the Browser carries
for regulation: open chromatin, histone marks, DNA methylation, transcription
factor binding, conservation and physical contact. It is adapted from Max's
draft, trimmed where it repeated the sections below it. The top of the page now
says when the track list was last checked, and CADD 1.7 goes ahead of
AlphaGenome among the variant scores, since it has been stable for years.
Six questions are new, each one chosen because people keep asking it on the
genome list: why a motif turns up across a whole gene, what a GeneHancer arc
does and does not claim, what the scores and grey shading mean (which differs
from track to track), whether signal heights can be compared (ENCODE4
auto-scales and ENCODE3 does not, which nothing else documents), what an empty
region means, and what to do when your tissue was never assayed.
The three "Which tracks show ..." headings become "How do I find ...", which is
the form the other FAQ pages use. The Single-cell ATAC-seq row leaves the
summary table because singleCellSignalsPeaks is still release alpha and those
links error anywhere but hgwdev; the row is kept in a comment to restore when
the track is released.
refs #24610
Co-Authored-By: Claude Opus 5 (1M context)
Return to FAQ Table of Contents
The assembly names after each track are links. They open that track's description page on
-that assembly, which gives the methods, the data version and the citation. Coverage varies a
+that assembly, which gives the methods, the data version, and the citation. Coverage varies a
lot between assemblies, so check the list before you assume a track exists on the genome you
work with.
+The tracks described here were last checked in October 2026. Projects such as JASPAR and
+ENCODE issue new versions on their own schedules, and we add tracks between checks, so use
+Track Search if you need to know what
+is on an assembly today.
The Genome Browser carries a large number of tracks that annotate regulatory regions. Most of
them are in the Regulation track group, which you will find below the browser image on
the main browser page.
+This page is about regulation at the level of DNA and chromatin. Mechanisms acting on RNA after
+it has been transcribed are a separate subject and are not covered. None of the tracks below
+observes regulation directly. Each one measures a property that regulatory regions tend to have,
+and the case for any particular region is built by stacking several of them. Almost all of these
+signals differ between cell types, so pick the tissue that matters for your question instead of
+reading a genome-wide summary.
A measured binding site comes from an experiment, usually ChIP-seq, in which one
protein was pulled down in one cell type under one set of conditions. The site is real in the
sense that the factor was found there in that experiment. It tells you nothing about other cell
types, and the experiment has to have been done for your factor and your tissue for the data to
exist at all.
A predicted binding site comes from scanning the genome sequence for a
motif, the sequence theme a given factor
prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist
everywhere in the genome for every
factor with a known motif, regardless of cell type, and most of them are not bound in
vivo. A typical transcription factor motif occurs hundreds of thousands of times in the
human genome, while the factor binds only a few thousand of those positions in any given cell.
Neither kind is better than the other. To find out where a factor was actually found, use a
measured track. To find out whether some sequence you care about, a variant or a promoter
fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a
prediction as evidence that the factor binds there.
For human, three tracks cover most needs. All three are in the Regulation group.
+This is the expected result for a motif search, not a sign that anything has gone wrong. Motifs
+are short and they tolerate mismatches, so a typical one occurs hundreds of thousands of times
+in the human genome, while the factor binds only a few thousand of those places in any given
+cell.
+Turning a matrix into a list of hits means picking a threshold, and there is no standard one.
+Two tools working from the same matrix will hand you different sites, because they cut at
+different places and weigh conservation differently. This has been an open problem for as long
+as people have been scanning genomes, so any particular set of predicted sites is best treated as
+one possible answer.
+Raising the score cutoff in the track settings will thin the display, but it does not make the
+survivors bound. A more productive approach is to read the motif as the list of places the
+factor could bind, then narrow that list using evidence that it did bind: a ChIP peak
+in a cell type close to yours, open chromatin, or both.
First check whether the experiment simply has not been done. ReMap covers the published ChIP-seq
experiments that were available when it was built, so if your factor is absent from ReMap there
may be no public ChIP-seq for it in that organism. The
ReMap website lets you search by target
and download the peaks per factor, and that is the quickest way to check.
If the experiment exists but is newer than our tracks, or was done in a cell type we do not
carry, you will need to load the data yourself as a custom track. The usual sources are:
If you would rather place a single file yourself, note that ENCODE distributes peaks as bigBed
and signal as bigWig, both of which the Browser reads directly. Copy the file URL from the portal;
you do not
need to download the file. Then paste one custom track line at
Add Custom Tracks:
Full instructions are on the
custom tracks help page and the
track hub help page.
-It depends on what you mean by a promoter, and the tracks disagree enough that it matters.Frequently Asked Questions: Regulation and cis-regulatory tracks
Topics
+
The basics
How do I tell whether a region is regulatory?
+
+
+
+
What is the difference between measured and predicted binding sites?
Transcription factor binding sites
-Which tracks show transcription factor binding sites?
+How do I find transcription factor binding sites?
+
+The JASPAR track shows my factor binding across the whole gene. Why?
+I cannot find my transcription factor in any track. Where else can I look?
track type=bigBed name="CTCF K562 peaks" bigDataUrl=https://www.encodeproject.org/files/ENCFF002CEL/@@download/ENCFF002CEL.bigBed
Promoters, enhancers and other elements
-Which tracks show promoters?
+How do I find the promoter of a gene?
For the chromatin evidence behind the called elements, see ENCODE4 Regulation, which carries DNase, ATAC-seq, histone modification and CTCF signal organized by tissue, on hg38 and mm10. This replaced ENCODE3 Regulation as the default in July 2026; ENCODE3 is kept for archival use and is still on hg19 and hg38.
+ ++It means that somebody has associated the two. It does not mean that the element has been +shown to regulate the gene.
++In GeneHancer the higher end of an arc is the element and the lower end is the +gene it has been linked to. The track pools several kinds of evidence to make those links. +They include eQTLs, promoter capture Hi-C, correlation between enhancer RNA and gene expression, +correlation between a transcription factor and a candidate target, and plain proximity. That +last one is worth remembering, because an arc drawn mostly on distance is making the same +nearest-gene assumption you were trying to get away from. The track's +double elite subset keeps only the cases where both the element and the link to the +gene were supported by more than one source. It is a great deal smaller and worth a look when +the full set gives you too much.
++The arcs also say nothing about the relationship between two genes that share an element. An +element linked to two genes has been associated with each of them separately, which is not a +statement that either gene affects the other.
+A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only and so on.
The ENCODE cCREs container on hg38 holds several versions. The ENCODE4 cCREs registry became the default in July 2026 and is the one to use: 2.3 million human and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with all four core assays. ENCODE3 cCREs is the earlier release, kept for archival use because a great many published analyses used it and coordinates need to stay reproducible. The container also carries per-biosample subtracks, which is how you restrict the classification to one cell type. Mouse mm10 carries the same three versions.
+ ++It often means signal strength, but not always, and the tracks on this page do not all use it +the same way. A score from 0 to 1000 drawn as a shade of grey is an old BED convention, dark for +high. Where a track uses that convention, the number reflects how strong the signal was in the +experiment behind the item. It is not a probability, and it is not a statement that the element +does anything.
++Beyond that convention, each track uses color for its own purpose. JASPAR does +use it as a score: the shade is the match to the motif profile, scaled to 0 to 1000, and sites +below 400 are hidden until the threshold is lowered in the track settings. +ReMap does not. It gives every transcription factor a color of its own, so the +color indicates which factor is shown and says nothing about the quality of the peak. +GeneHancer uses color for two things at once, whether the element is a +promoter or an enhancer and whether its confidence is high, medium or low. The +cCRE tracks color by class, promoter-like against enhancer-like against +CTCF-only, which is unrelated to strength.
++The track description page is authoritative in each case, and it is one click from the track +name.
+ + ++Between cell types in one window, yes. That is what the layered tracks are built for: the +colored traces share a vertical axis, so a taller trace really is a stronger signal.
++Between regions, check the scale first. The ENCODE4 histone and accessibility +tracks on hg38 arrive set to +Auto-scale to data view, which recomputes the vertical range from whatever is on screen. +The same peak is then drawn at different heights depending on how far out you are zoomed. Two +screenshots of two places cannot be compared at all. Open the +track's configuration page, change Data view scaling to use the vertical viewing range, +and set a range that suits your data. The older ENCODE3 layered tracks on +hg19 and +hg38 come with a fixed +range instead, so they are already comparable from one place to the next.
++Between one track and another, no. The experiments were done by different labs and run through +different pipelines, so the numbers are not in the same units. A value of 12 in one track and a +value of 12 in another are not the same quantity, and the display gives no warning of this.
++When the comparison matters, it is better to work with the values than with the picture. The +Table Browser and the +Data Integrator will return values for your regions, and +the track description page says what the units are.
+ + ++No. It means these particular experiments did not report anything in that region. +
++The most common reason is that nobody has assayed your cell type. The tracks we +carry cover only a selective number of cell types or tissues chosen by the +experimenters, and a region may be a strong enhancer in a tissue that has never +been profiled. The same goes for the factor. ReMap and the ENCODE ChIP tracks +together cover a few hundred factors in a few hundred cell types, which is a +small fraction of the possible combinations. Peak calling also requires a +cutoff, so a site that is real but weak can fall below it and thus will not +be annotated.
++Annotations that do not depend on a particular experiment make a better test for absence. +Conservation requires no experiment at all, and the cCREs pool a great many biosamples into one +classification. If those are also empty over the region, there is more reason to think it is +may not be a regulatory region.
+-Most of the large regulatory tracks are collections holding many separate tracks, one per cell -type or experiment, and they arrive with only a summary view turned on. Click the track name to +Most of the large regulatory tracks are collections holding many separate tracks, often with +one per cell type or experiment. By default, these collections often only have a summary view +turned on. Click the track name to open its configuration page, where you will find the list and, on the bigger collections, filters. ReMap, for instance, lets you filter by transcription factor directly in the track settings, so you can show one factor across all its experiments.
The largest collections use a different configuration page. Where there are hundreds or thousands of individual experiments, as in the per-experiment tracks of ENCODE4 Regulation on hg38, the page shows facets down the left side and a paginated table on the right: tick the tissue, assay or biosample you want and the table narrows to the matching experiments. This is usually a faster way to reach one cell type than reading a long list.
To find a track rather than configure one, the Track Search page searches track names and descriptions across the whole assembly, which is usually faster than reading through the track groups.
+ ++One solution may be to choose the closest available cell type, and state that choice in any +write-up. For example, treating a cell line as a stand-in for primary cells is a judgment made by +the person reading the data, not something the data supports on its own.
++Before you settle for a substitute, check whether the experiment exists somewhere we do not +carry it. The ENCODE portal holds +far more biosamples than we turn into tracks. It will open any of them in the Browser for you; +see displaying ENCODE data above. You can also search the +public hubs list, as some may contain more +tissue-specific regulatory data than what's available natively in the Browser.
++If nothing close exists, consider the annotations that were never specific to one cell type. +The cCREs pool many biosamples into a single classification, and conservation does not depend on +any experiment. Neither one describes your tissue. Neither one silently describes a different +tissue either, which is the risk that comes with a poorly matched proxy.
+No single track answers this, so you have to work outward from the gene.
MPRAVarDB holds 239,028 variants that were put through a reporter assay and scored for allelic effect, drawn from 18 MPRA studies covering more than 30 cell lines and more than 30 diseases or traits. The variants come from GWAS and eQTL fine-mapping, from saturation mutagenesis of 20 disease-associated regulatory elements, and from smaller focused screens, so the coverage is concentrated on loci people have already had reason to care about. Items are colored by significance, dark red for FDR below 0.05. If your variant is in here, you have an experimental answer and not a guess. It lives in the MPRAs collection, alongside MPRA Base, which tests whole elements. Available for hg38.
Most variants will not be in it, since 239,028 tested positions is a very small part of the -genome. Two prediction tracks fill the gap by scoring every possible substitution, and neither -is in the Regulation group: look under Phenotype and Disease Associations, in -the Deleteriousness Predictions collection on +genome. Three prediction tracks fill the gap by scoring every possible substitution. None of +them is in the Regulation group; all three are under Phenotype and Disease +Associations, the last two inside the Deleteriousness Predictions +collection on hg38.
A high score from either says a model thinks the change is disruptive, not that anyone has measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement is the better evidence.
Much less. The limit is usually the data itself: these resources were built for human first, and @@ -511,59 +740,70 @@