b3aef6de6abdf2356bf7d4a64fd5d857726eea6b mspeir Sun Oct 4 09:57:49 2026 -0700 Regulation FAQ: cis-regulation background and six interpretive questions Max asked for a short introduction to cis-regulation, and Lou noted that the page leaned towards listing datasets rather than explaining how to read them. This covers both. A new opening question walks through the kinds of evidence the Browser carries for regulation: open chromatin, histone marks, DNA methylation, transcription factor binding, conservation and physical contact. It is adapted from Max's draft, trimmed where it repeated the sections below it. The top of the page now says when the track list was last checked, and CADD 1.7 goes ahead of AlphaGenome among the variant scores, since it has been stable for years. Six questions are new, each one chosen because people keep asking it on the genome list: why a motif turns up across a whole gene, what a GeneHancer arc does and does not claim, what the scores and grey shading mean (which differs from track to track), whether signal heights can be compared (ENCODE4 auto-scales and ENCODE3 does not, which nothing else documents), what an empty region means, and what to do when your tissue was never assayed. The three "Which tracks show ..." headings become "How do I find ...", which is the form the other FAQ pages use. The Single-cell ATAC-seq row leaves the summary table because singleCellSignalsPeaks is still release alpha and those links error anywhere but hgwdev; the row is kept in a comment to restore when the track is released. refs #24610 Co-Authored-By: Claude Opus 5 (1M context) diff --git src/hg/htdocs/FAQ/FAQregulation.html src/hg/htdocs/FAQ/FAQregulation.html index 8e06956e0c6..64febc9776f 100755 --- src/hg/htdocs/FAQ/FAQregulation.html +++ src/hg/htdocs/FAQ/FAQregulation.html @@ -1,577 +1,817 @@

Frequently Asked Questions: Regulation and cis-regulatory tracks

Topics


Return to FAQ Table of Contents

The assembly names after each track are links. They open that track's description page on -that assembly, which gives the methods, the data version and the citation. Coverage varies a +that assembly, which gives the methods, the data version, and the citation. Coverage varies a lot between assemblies, so check the list before you assume a track exists on the genome you work with.

+

+The tracks described here were last checked in October 2026. Projects such as JASPAR and +ENCODE issue new versions on their own schedules, and we add tracks between checks, so use +Track Search if you need to know what +is on an assembly today.

- +

The basics

The Genome Browser carries a large number of tracks that annotate regulatory regions. Most of them are in the Regulation track group, which you will find below the browser image on the main browser page.

+
How do I tell whether a region is regulatory?
+

+This page is about regulation at the level of DNA and chromatin. Mechanisms acting on RNA after +it has been transcribed are a separate subject and are not covered. None of the tracks below +observes regulation directly. Each one measures a property that regulatory regions tend to have, +and the case for any particular region is built by stacking several of them. Almost all of these +signals differ between cell types, so pick the tissue that matters for your question instead of +reading a genome-wide summary.

+ + +
What is the difference between measured and predicted binding sites?

A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.

A predicted binding site comes from scanning the genome sequence for a motif, the sequence theme a given factor prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.

Neither kind is better than the other. To find out where a factor was actually found, use a measured track. To find out whether some sequence you care about, a variant or a promoter fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a prediction as evidence that the factor binds there.

Transcription factor binding sites

-
Which tracks show transcription factor binding sites?
+
How do I find transcription factor binding sites?

For human, three tracks cover most needs. All three are in the Regulation group.

+ +
The JASPAR track shows my factor binding across the whole gene. Why?
+

+This is the expected result for a motif search, not a sign that anything has gone wrong. Motifs +are short and they tolerate mismatches, so a typical one occurs hundreds of thousands of times +in the human genome, while the factor binds only a few thousand of those places in any given +cell.

+

+Turning a matrix into a list of hits means picking a threshold, and there is no standard one. +Two tools working from the same matrix will hand you different sites, because they cut at +different places and weigh conservation differently. This has been an open problem for as long +as people have been scanning genomes, so any particular set of predicted sites is best treated as +one possible answer.

+

+Raising the score cutoff in the track settings will thin the display, but it does not make the +survivors bound. A more productive approach is to read the motif as the list of places the +factor could bind, then narrow that list using evidence that it did bind: a ChIP peak +in a cell type close to yours, open chromatin, or both.

+
I cannot find my transcription factor in any track. Where else can I look?

First check whether the experiment simply has not been done. ReMap covers the published ChIP-seq experiments that were available when it was built, so if your factor is absent from ReMap there may be no public ChIP-seq for it in that organism. The ReMap website lets you search by target and download the peaks per factor, and that is the quickest way to check.

If the experiment exists but is newer than our tracks, or was done in a cell type we do not carry, you will need to load the data yourself as a custom track. The usual sources are:

How do I display ENCODE data that is not already a track?

You do not need to download anything or write a custom track by hand. The ENCODE portal will open its data in the Genome Browser for you.

  1. Search the ENCODE portal for what you want, narrowing the results with the filters down the left side. Assay title, Target of assay (the factor), Biosample (the cell type or tissue) and Genome assembly are the useful ones.
  2. Click the Visualize button above the result list.
  3. Pick your assembly in the panel that opens, then click UCSC.

The Browser opens with every experiment in your filtered result set loaded as a track hub, so this works just as well for one experiment as for fifty. Restrict the search before you visualize, since a broad filter can attach a very large number of tracks at once.

ENCODE also publishes a hub for each individual experiment, which is handy if you are scripting or want to keep a link in a session. Substitute the accession into this URL:

https://www.encodeproject.org/experiments/ENCSR000AKO/@@hub/hub.txt

and load it from the My Hubs tab of the Track Hubs page, or by appending it to a browser URL as hgTracks?db=hg38&hubUrl= followed by the hub address.

Promoters, enhancers and other elements

-
Which tracks show promoters?
+
How do I find the promoter of a gene?

-It depends on what you mean by a promoter, and the tracks disagree enough that it matters.

+It depends on what you mean by a promoter, something which the tracks don't always agree on. In +the loosest sense a promoter is the minimal stretch upstream of the transcription start site +that will drive expression, but there is no agreed figure for minimal: some authors work with +2 kb upstream, others with 500 bp or 100 bp, and the right choice depends on the gene and on +whether you are after broad or tissue-specific expression.

-
Which tracks show enhancers and other regulatory elements?
+
How do I find enhancers and other regulatory elements?

For the chromatin evidence behind the called elements, see ENCODE4 Regulation, which carries DNase, ATAC-seq, histone modification and CTCF signal organized by tissue, on hg38 and mm10. This replaced ENCODE3 Regulation as the default in July 2026; ENCODE3 is kept for archival use and is still on hg19 and hg38.

+ +
What does it mean when GeneHancer draws an arc between an element and a gene?
+

+It means that somebody has associated the two. It does not mean that the element has been +shown to regulate the gene.

+

+In GeneHancer the higher end of an arc is the element and the lower end is the +gene it has been linked to. The track pools several kinds of evidence to make those links. +They include eQTLs, promoter capture Hi-C, correlation between enhancer RNA and gene expression, +correlation between a transcription factor and a candidate target, and plain proximity. That +last one is worth remembering, because an arc drawn mostly on distance is making the same +nearest-gene assumption you were trying to get away from. The track's +double elite subset keeps only the cases where both the element and the link to the +gene were supported by more than one source. It is a great deal smaller and worth a look when +the full set gives you too much.

+

+The arcs also say nothing about the relationship between two genes that share an element. An +element linked to two genes has been associated with each of them separately, which is not a +statement that either gene affects the other.

+
What are cCREs, and which cCRE track should I use?

A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only and so on.

The ENCODE cCREs container on hg38 holds several versions. The ENCODE4 cCREs registry became the default in July 2026 and is the one to use: 2.3 million human and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with all four core assays. ENCODE3 cCREs is the earlier release, kept for archival use because a great many published analyses used it and coordinates need to stay reproducible. The container also carries per-biosample subtracks, which is how you restrict the classification to one cell type. Mouse mm10 carries the same three versions.

+ +

Reading the data

+ +
What do the scores and the grey shading in these tracks mean?
+

+It often means signal strength, but not always, and the tracks on this page do not all use it +the same way. A score from 0 to 1000 drawn as a shade of grey is an old BED convention, dark for +high. Where a track uses that convention, the number reflects how strong the signal was in the +experiment behind the item. It is not a probability, and it is not a statement that the element +does anything.

+

+Beyond that convention, each track uses color for its own purpose. JASPAR does +use it as a score: the shade is the match to the motif profile, scaled to 0 to 1000, and sites +below 400 are hidden until the threshold is lowered in the track settings. +ReMap does not. It gives every transcription factor a color of its own, so the +color indicates which factor is shown and says nothing about the quality of the peak. +GeneHancer uses color for two things at once, whether the element is a +promoter or an enhancer and whether its confidence is high, medium or low. The +cCRE tracks color by class, promoter-like against enhancer-like against +CTCF-only, which is unrelated to strength.

+

+The track description page is authoritative in each case, and it is one click from the track +name.

+ + +
The signal is higher in one region than another. Can I compare them?
+

+Between cell types in one window, yes. That is what the layered tracks are built for: the +colored traces share a vertical axis, so a taller trace really is a stronger signal.

+

+Between regions, check the scale first. The ENCODE4 histone and accessibility +tracks on hg38 arrive set to +Auto-scale to data view, which recomputes the vertical range from whatever is on screen. +The same peak is then drawn at different heights depending on how far out you are zoomed. Two +screenshots of two places cannot be compared at all. Open the +track's configuration page, change Data view scaling to use the vertical viewing range, +and set a range that suits your data. The older ENCODE3 layered tracks on +hg19 and +hg38 come with a fixed +range instead, so they are already comparable from one place to the next.

+

+Between one track and another, no. The experiments were done by different labs and run through +different pipelines, so the numbers are not in the same units. A value of 12 in one track and a +value of 12 in another are not the same quantity, and the display gives no warning of this.

+

+When the comparison matters, it is better to work with the values than with the picture. The +Table Browser and the +Data Integrator will return values for your regions, and +the track description page says what the units are.

+ + +
Nothing is annotated over my region. Does that mean it is not regulatory?
+

+No. It means these particular experiments did not report anything in that region. +

+

+The most common reason is that nobody has assayed your cell type. The tracks we +carry cover only a selective number of cell types or tissues chosen by the +experimenters, and a region may be a strong enhancer in a tissue that has never +been profiled. The same goes for the factor. ReMap and the ENCODE ChIP tracks +together cover a few hundred factors in a few hundred cell types, which is a +small fraction of the possible combinations. Peak calling also requires a +cutoff, so a site that is real but weak can fall below it and thus will not +be annotated.

+

+Annotations that do not depend on a particular experiment make a better test for absence. +Conservation requires no experiment at all, and the cCREs pool a great many biosamples into one +classification. If those are also empty over the region, there is more reason to think it is +may not be a regulatory region.

+

Working with the data

How do I restrict a search to one cell type or tissue?

-Most of the large regulatory tracks are collections holding many separate tracks, one per cell -type or experiment, and they arrive with only a summary view turned on. Click the track name to +Most of the large regulatory tracks are collections holding many separate tracks, often with +one per cell type or experiment. By default, these collections often only have a summary view +turned on. Click the track name to open its configuration page, where you will find the list and, on the bigger collections, filters. ReMap, for instance, lets you filter by transcription factor directly in the track settings, so you can show one factor across all its experiments.

The largest collections use a different configuration page. Where there are hundreds or thousands of individual experiments, as in the per-experiment tracks of ENCODE4 Regulation on hg38, the page shows facets down the left side and a paginated table on the right: tick the tissue, assay or biosample you want and the table narrows to the matching experiments. This is usually a faster way to reach one cell type than reading a long list.

To find a track rather than configure one, the Track Search page searches track names and descriptions across the whole assembly, which is usually faster than reading through the track groups.

+ +
My tissue is not in the list of cell types. What should I do?
+

+One solution may be to choose the closest available cell type, and state that choice in any +write-up. For example, treating a cell line as a stand-in for primary cells is a judgment made by +the person reading the data, not something the data supports on its own.

+

+Before you settle for a substitute, check whether the experiment exists somewhere we do not +carry it. The ENCODE portal holds +far more biosamples than we turn into tracks. It will open any of them in the Browser for you; +see displaying ENCODE data above. You can also search the +public hubs list, as some may contain more +tissue-specific regulatory data than what's available natively in the Browser.

+

+If nothing close exists, consider the annotations that were never specific to one cell type. +The cCREs pool many biosamples into a single classification, and conservation does not depend on +any experiment. Neither one describes your tissue. Neither one silently describes a different +tissue either, which is the risk that comes with a poorly matched proxy.

+
I have a gene. How do I find the factors that regulate it?

No single track answers this, so you have to work outward from the gene.

  1. Navigate to the gene and zoom out far enough to include the surrounding non-coding sequence. Regulatory elements are often tens or hundreds of kilobases away, and the nearest gene to an element is frequently not its target.
  2. Turn on GeneHancer. Its interaction arcs will show which elements have been linked to your gene, including distant ones, which narrows the search from the whole neighborhood to a handful of regions.
  3. Turn on ReMap or TF ChIP and look at which factors have peaks in those regions. This gives you factors that were measured at that position in some cell type.
  4. Check whether any of those cell types are relevant to your biology. A peak in K562 says little about neurons.
  5. If you need candidates in a cell type nobody has assayed, fall back to JASPAR predictions within the GeneHancer elements, and treat the result as hypotheses to test.

To do this systematically, the Data Integrator will intersect two or more tracks and return a table, and the Table Browser will do the same for a region or for a list of genes.

I have a variant, not a region. What does it do to regulation?

This is a different question from the rest of this page. Everything above annotates regions and tells you what is where; the tracks here take one base change and tell you what it does. The same measured and predicted distinction applies, so start with the one track that is measured.

MPRAVarDB holds 239,028 variants that were put through a reporter assay and scored for allelic effect, drawn from 18 MPRA studies covering more than 30 cell lines and more than 30 diseases or traits. The variants come from GWAS and eQTL fine-mapping, from saturation mutagenesis of 20 disease-associated regulatory elements, and from smaller focused screens, so the coverage is concentrated on loci people have already had reason to care about. Items are colored by significance, dark red for FDR below 0.05. If your variant is in here, you have an experimental answer and not a guess. It lives in the MPRAs collection, alongside MPRA Base, which tests whole elements. Available for hg38.

Most variants will not be in it, since 239,028 tested positions is a very small part of the -genome. Two prediction tracks fill the gap by scoring every possible substitution, and neither -is in the Regulation group: look under Phenotype and Disease Associations, in -the Deleteriousness Predictions collection on +genome. Three prediction tracks fill the gap by scoring every possible substitution. None of +them is in the Regulation group; all three are under Phenotype and Disease +Associations, the last two inside the Deleteriousness Predictions +collection on hg38.

A high score from either says a model thinks the change is disruptive, not that anyone has measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement is the better evidence.

What is available for assemblies other than human and mouse?

Much less. The limit is usually the data itself: these resources were built for human first, and most never went further than mouse. JASPAR is the main exception, since it needs only the genome sequence and a motif, so it covers zebrafish, fly, worm, chicken, sea squirt and yeast as well. ReMap also covers fly. Everything else described on this page is human and mouse only.

Mouse is split awkwardly between its own assemblies. mm10 carries most of these tracks and mm39 has only JASPAR, ReMap and VISTA, because several of the source projects never released mm39 versions. If you need something that is on mm10 but not mm39, LiftOver will convert the coordinates, though check the result before trusting it. Human has a milder version of the same problem: a few tracks, such as the older clustered Txn Factor ChIP, are still hg19 only.

For assemblies not hosted at UCSC, or for tracks we do not carry, check the public hubs list, where other groups publish data through our browser.

Summary: regulatory tracks by category

Each assembly name below links to that track's description page on that assembly. A few tracks appear on additional genomes not listed here; use Track Search to check a genome that is not shown.

+ + + + + + +
Track What it is Measured or predicted Assemblies
Transcription factor binding
ReMap ChIP-seq Public ChIP-seq for transcriptional regulators, integrated Measured hg19, hg38, mm10, mm39, dm6
TF ChIP (ENCODE 3 TFBS on hg19) ENCODE 3 TF ChIP-seq peaks, around 340 factors in about 130 cell types Measured hg19, hg38
JASPAR Transcription Factors Motif matches from the JASPAR CORE collection; several releases as subtracks Predicted JASPAR 2026 on hg38, mm39, danRer11, galGal6, dm6, ce11, ci3, sacCer3; up to JASPAR 2024 on hg19 and mm10
Promoters and transcription start sites
EPDnew Promoters Experimentally defined promoters with mapped start sites Measured hg19, hg38, mm10
FANTOM5 CAGE transcription start sites and their usage per tissue Measured hg19, hg38, mm10
Enhancers and candidate elements
GeneHancer Regulatory elements linked to predicted target genes Mixed, with predicted targets hg19, hg38
VISTA Enhancers Elements tested individually in transgenic mouse assays Measured, validated hg19, hg38, mm10, mm39
ENCODE cCREs Candidate elements classified from chromatin signal; ENCODE4 is the default Predicted from measured signal hg38, mm10
RefSeq Functional Elements NCBI curated non-coding functional elements Measured, curated hg38, mm10
MPRAs (MPRA Base) Reporter assay activity for 40,938 tested regulatory elements Measured hg38
Chromatin and expression context
ENCODE4 Regulation DNase, ATAC, histone marks and CTCF by tissue; the current default Measured hg38, mm10
ENCODE3 Regulation DNase, histone marks and transcription signal; archival, replaced by ENCODE4 Measured hg19, hg38
GTEx Gene Gene expression across 53 tissues Measured hg19, hg38
GTEx cis-eQTLs Variants associated with expression of nearby genes Measured hg38
Effect of a single variant
MPRAVarDB 239,028 variants tested for allelic effect in reporter assays Measured hg38
CADD 1.7Deleteriousness score for every single-base substitution, coding and non-coding; + under Phenotype and Disease AssociationsPredictedhg19, + hg38
AlphaGenome Variant Impact score for every single-base substitution, coding and non-coding; under Phenotype and Disease Associations Predicted hg38
PromoterAI Score for every single-base substitution in proximal promoters; under Phenotype and Disease Associations Predicted hg38

For the full set of tracks on any assembly, open the Regulation group on the browser page, or use Track Search.