4fa04b06c8c9e1627bb6bedef9e78a02fc4ceb60 mspeir Sat Sep 26 16:25:41 2026 -0700 minor tewaks to wording; removing oreganno refs; removing params from hgTracks links so they use users cart, refs #24610 diff --git src/hg/htdocs/FAQ/FAQregulation.html src/hg/htdocs/FAQ/FAQregulation.html index 611ee9f2595..22999c560f5 100755 --- src/hg/htdocs/FAQ/FAQregulation.html +++ src/hg/htdocs/FAQ/FAQregulation.html @@ -1,612 +1,575 @@

Frequently Asked Questions: Regulation and cis-regulatory tracks

Topics


Return to FAQ Table of Contents

The assembly names after each track are links. They open that track's description page on that assembly, which gives the methods, the data version and the citation. Coverage varies a lot between assemblies, so check the list before you assume a track exists on the genome you work with.

The basics

The Genome Browser carries a large number of tracks that annotate regulatory regions. Most of them are in the Regulation track group, which you will find below the browser image on -the main browser page. It is not always obvious which -one you want, because these tracks answer quite different questions even when they all look -like boxes on the screen.

+the main browser page.

What is the difference between measured and predicted binding sites?

-Nearly every question we get about these tracks comes down to this distinction.

-

A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.

A predicted binding site comes from scanning the genome sequence for a motif, the sequence theme a given factor prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.

Neither kind is better than the other. To find out where a factor was actually found, use a measured track. To find out whether some sequence you care about, a variant or a promoter fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a prediction as evidence that the factor binds there.

Transcription factor binding sites

Which tracks show transcription factor binding sites?

For human, three tracks cover most needs. All three are in the Regulation group.

-

-Two older tracks still come up in questions. TFBS Conserved shows predicted -sites that are conserved across human, mouse and rat, on -hg17, -hg18 and -hg19 only; it has not been updated in -many years and -there is no hg38 version. ORegAnno is a curated collection of regulatory -elements taken from the literature, so it is small but every entry has a citation, on -hg19, -hg38, -mm10 and -dm6.

I cannot find my transcription factor in any track. Where else can I look?

First check whether the experiment simply has not been done. ReMap covers the published ChIP-seq experiments that were available when it was built, so if your factor is absent from ReMap there may be no public ChIP-seq for it in that organism. The ReMap website lets you search by target and download the peaks per factor, and that is the quickest way to check.

If the experiment exists but is newer than our tracks, or was done in a cell type we do not -carry, you will need to bring the data in yourself. The usual sources are:

+carry, you will need to load the data yourself as a custom track. The usual sources are:

How do I display ENCODE data that is not already a track?

You do not need to download anything or write a custom track by hand. The ENCODE portal will open its data in the Genome Browser for you.

  1. Search the ENCODE portal for what you want, narrowing the results with the filters down the left side. Assay title, Target of assay (the factor), Biosample (the cell type or tissue) and Genome assembly are the useful ones.
  2. Click the Visualize button above the result list.
  3. Pick your assembly in the panel that opens, then click UCSC.

The Browser opens with every experiment in your filtered result set loaded as a track hub, so this works just as well for one experiment as for fifty. Restrict the search before you visualize, since a broad filter can attach a very large number of tracks at once.

ENCODE also publishes a hub for each individual experiment, which is handy if you are scripting or want to keep a link in a session. Substitute the accession into this URL:

https://www.encodeproject.org/experiments/ENCSR000AKO/@@hub/hub.txt

and load it from the My Hubs tab of the Track Hubs page, or by appending it to a browser URL as hgTracks?db=hg38&hubUrl= followed by the hub address.

If you would rather place a single file yourself, note that ENCODE distributes peaks as bigBed and signal as bigWig, both of which the Browser reads directly. Copy the file URL from the portal; you do not need to download the file. Then paste one custom track line at Add Custom Tracks:

track type=bigBed name="CTCF K562 peaks" bigDataUrl=https://www.encodeproject.org/files/ENCFF002CEL/@@download/ENCFF002CEL.bigBed

-The Browser does not download the whole file. bigBed and bigWig are indexed, so it fetches only -the part covering the region you are looking at, which is why a multi-gigabyte signal file opens -in a moment. Full instructions are on the +Full instructions are on the custom tracks help page and the track hub help page.

Promoters, enhancers and other elements

Which tracks show promoters?

It depends on what you mean by a promoter, and the tracks disagree enough that it matters.

Which tracks show enhancers and other regulatory elements?

For the chromatin evidence behind the called elements, see ENCODE4 Regulation, which carries DNase, ATAC-seq, histone modification and CTCF signal organized by tissue, on hg38 and mm10. This replaced ENCODE3 Regulation as the default in July 2026; ENCODE3 is kept for archival use and is still on hg19 and hg38.

What are cCREs, and which cCRE track should I use?

A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only -and so on. The word candidate is doing real work here: these are regions worth testing, -not confirmed regulatory elements.

+and so on.

The ENCODE cCREs container on hg38 holds several versions. The ENCODE4 cCREs registry became the default in July 2026 and is the one to use: 2.3 million human and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with all four core assays. ENCODE3 cCREs is the earlier release, kept for archival use because a great many published analyses used it and coordinates need to stay reproducible. The container also carries per-biosample subtracks, which is how you restrict the classification to one cell type. Mouse mm10 carries the same three versions.

Working with the data

How do I restrict a search to one cell type or tissue?

Most of the large regulatory tracks are collections holding many separate tracks, one per cell type or experiment, and they arrive with only a summary view turned on. Click the track name to open its configuration page, where you will find the list and, on the bigger collections, filters. ReMap, for instance, lets you filter by transcription factor directly in the track settings, so you can show one factor across all its experiments.

The largest collections use a different configuration page. Where there are hundreds or thousands of individual experiments, as in the per-experiment tracks of ENCODE4 Regulation on hg38, the page shows facets down the left side and a paginated table on the right: tick the tissue, assay or biosample you want and the table narrows to the matching experiments. This is usually a faster way to reach one cell type than reading a long list.

To find a track rather than configure one, the Track Search page searches track names and descriptions across the whole assembly, which is usually faster than reading through the track groups.

I have a gene. How do I find the factors that regulate it?

No single track answers this, so you have to work outward from the gene.

  1. Navigate to the gene and zoom out far enough to include the surrounding non-coding sequence. Regulatory elements are often tens or hundreds of kilobases away, and the nearest gene to an element is frequently not its target.
  2. Turn on GeneHancer. Its interaction arcs will show which elements have been linked to your gene, including distant ones, which narrows the search from the whole neighborhood to a handful of regions.
  3. Turn on ReMap or TF ChIP and look at which factors have peaks in those regions. This gives you factors that were measured at that position in some cell type.
  4. Check whether any of those cell types are relevant to your biology. A peak in K562 says little about neurons.
  5. If you need candidates in a cell type nobody has assayed, fall back to JASPAR predictions within the GeneHancer elements, and treat the result as hypotheses to test.

To do this systematically, the Data Integrator will intersect two or more tracks and return a table, and the Table Browser will do the same for a region or for a list of genes.

I have a variant, not a region. What does it do to regulation?

This is a different question from the rest of this page. Everything above annotates regions and tells you what is where; the tracks here take one base change and tell you what it does. The same measured and predicted distinction applies, so start with the one track that is measured.

MPRAVarDB holds 239,028 variants that were put through a reporter assay and scored for allelic effect, drawn from 18 MPRA studies covering more than 30 cell lines and more than 30 diseases or traits. The variants come from GWAS and eQTL fine-mapping, from saturation mutagenesis of 20 disease-associated regulatory elements, and from smaller focused screens, so the coverage is concentrated on loci people have already had reason to care about. Items are colored by significance, dark red for FDR below 0.05. If your variant is in here, you have an experimental answer and not a guess. It lives in the MPRAs collection, alongside MPRA Base, which tests whole elements. Available for hg38.

Most variants will not be in it, since 239,028 tested positions is a very small part of the genome. Two prediction tracks fill the gap by scoring every possible substitution, and neither is in the Regulation group: look under Phenotype and Disease Associations, in the Deleteriousness Predictions collection on hg38.

A high score from either says a model thinks the change is disruptive, not that anyone has measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement is the better evidence.

What is available for assemblies other than human and mouse?

Much less. The limit is usually the data itself: these resources were built for human first, and most never went further than mouse. JASPAR is the main exception, since it needs only the genome sequence and a motif, so it covers zebrafish, fly, worm, chicken, sea squirt and yeast as well. -ReMap and ORegAnno also cover fly. Everything else described on this page is human and mouse +ReMap also covers fly. Everything else described on this page is human and mouse only.

Mouse is split awkwardly between its own assemblies. mm10 carries most of these tracks and mm39 has only JASPAR, ReMap and VISTA, because several of the source projects never released mm39 versions. If you need something that is on mm10 but not mm39, LiftOver will convert the coordinates, though check the -result before trusting it. Human has a milder version of the same problem: a few tracks are -still hg19 only, and TFBS Conserved was never rebuilt for hg38.

+result before trusting it. Human has a milder version of the same problem: a few tracks, such +as the older clustered Txn Factor ChIP, are still hg19 only.

For assemblies not hosted at UCSC, or for tracks we do not carry, check the public hubs list, where other groups publish data through our browser.

Summary: regulatory tracks by category

Each assembly name below links to that track's description page on that assembly. A few tracks appear on additional genomes not listed here; use Track Search to check a genome that is not shown.

- - - - - - - - - - - -
Track What it is Measured or predicted Assemblies
Transcription factor binding
ReMap ChIP-seq Public ChIP-seq for transcriptional regulators, integrated Measured hg19, hg38, mm10, mm39, dm6
TF ChIP (ENCODE 3 TFBS on hg19) ENCODE 3 TF ChIP-seq peaks, around 340 factors in about 130 cell types Measured hg19, hg38
JASPAR Transcription Factors Motif matches from the JASPAR CORE collection; several releases as subtracks Predicted JASPAR 2026 on hg38, mm39, danRer11, galGal6, dm6, ce11, ci3, sacCer3; up to JASPAR 2024 on hg19 and mm10
TFBS ConservedConserved motif matches; not updated recently, no hg38 versionPredictedhg17, - hg18, - hg19
ORegAnnoRegulatory elements curated from the literatureMeasured, curatedhg19, - hg38, - mm10, - dm6
Promoters and transcription start sites
EPDnew Promoters Experimentally defined promoters with mapped start sites Measured hg19, hg38, mm10
FANTOM5 CAGE transcription start sites and their usage per tissue Measured hg19, hg38, mm10
Enhancers and candidate elements
GeneHancer Regulatory elements linked to predicted target genes Mixed, with predicted targets hg19, hg38
VISTA Enhancers Elements tested individually in transgenic mouse assays Measured, validated hg19, hg38, mm10, mm39
ENCODE cCREs Candidate elements classified from chromatin signal; ENCODE4 is the default Predicted from measured signal hg38, mm10
RefSeq Functional Elements NCBI curated non-coding functional elements Measured, curated hg38, mm10
MPRAs (MPRA Base) Reporter assay activity for 40,938 tested regulatory elements Measured hg38
Chromatin and expression context
ENCODE4 Regulation DNase, ATAC, histone marks and CTCF by tissue; the current default Measured hg38, mm10
ENCODE3 Regulation DNase, histone marks and transcription signal; archival, replaced by ENCODE4 Measured hg19, hg38
Single-cell ATAC-seq Accessibility peaks and signal from Cell Browser datasets Measured hg38, mm10
GTEx Gene Gene expression across 53 tissues Measured hg19, hg38
GTEx cis-eQTLs Variants associated with expression of nearby genes Measured hg38
Effect of a single variant
MPRAVarDB 239,028 variants tested for allelic effect in reporter assays Measured hg38
AlphaGenome Variant Impact score for every single-base substitution, coding and non-coding; under Phenotype and Disease Associations Predicted hg38
PromoterAI Score for every single-base substitution in proximal promoters; under Phenotype and Disease Associations Predicted hg38

For the full set of tracks on any assembly, open the Regulation group on the -browser page, or use +browser page, or use Track Search.