5fd99d465a31d5bd38320ac85ff1969ab9515014 mspeir Wed Sep 23 15:44:06 2026 -0700 Regulation FAQ: correct assembly coverage, add variant effect section, refs #24610 Corrections found by checking the page against our own release announcements, which I should have done before the first commit: - ENCODE4 Regulation is on mm10 as well as hg38. The track symbol differs by assembly (wgEncodeReg4 on hg38, encode4Reg on mm10), so a search by name missed it. - JASPAR 2026 covers hg38, mm39, danRer11, galGal6, dm6, ce11, ci3 and sacCer3. hg19 and mm10 stop at JASPAR 2024. The page had implied all ten were 2026. Also notes that the track holds several releases as subtracks, so the version depends on which one is turned on. - TFBS Conserved is on hg17 as well as hg18 and hg19. - ENCODE4 replaced ENCODE3 as the default in July 2026 and ENCODE3 is kept for archival use. Said in the text and marked in the summary table. - The cell type section described the old subtrack list for the large collections, which now use the faceted interface. - Dropped "composite" and "superTrack", which users should not see. New section on what a single variant does to regulation, covering MPRAVarDB for variants that were actually tested in a reporter assay, and AlphaGenome and PromoterAI for predictions. These live under Phenotype and Disease Associations rather than Regulation, which is worth saying since that is where a reader would otherwise fail to find them. Links to goldenPath/help/hgRegMotifHelp.html from the motif definition and from the JASPAR entry, and the definition now matches the wording there. Co-Authored-By: Claude Opus 5 (1M context) diff --git src/hg/htdocs/FAQ/FAQregulation.html src/hg/htdocs/FAQ/FAQregulation.html index 575f02c1a79..611ee9f2595 100755 --- src/hg/htdocs/FAQ/FAQregulation.html +++ src/hg/htdocs/FAQ/FAQregulation.html @@ -1,521 +1,612 @@

Frequently Asked Questions: Regulation and cis-regulatory tracks

Topics


Return to FAQ Table of Contents

The assembly names after each track are links. They open that track's description page on that assembly, which gives the methods, the data version and the citation. Coverage varies a lot between assemblies, so check the list before you assume a track exists on the genome you work with.

The basics

The Genome Browser carries a large number of tracks that annotate regulatory regions. Most of them are in the Regulation track group, which you will find below the browser image on the main browser page. It is not always obvious which one you want, because these tracks answer quite different questions even when they all look like boxes on the screen.

What is the difference between measured and predicted binding sites?

Nearly every question we get about these tracks comes down to this distinction.

A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.

-A predicted binding site comes from scanning the genome sequence for a motif, a short -pattern that the factor is known to prefer. Predictions exist everywhere in the genome for every +A predicted binding site comes from scanning the genome sequence for a +motif, the sequence theme a given factor +prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist +everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.

Neither kind is better than the other. To find out where a factor was actually found, use a measured track. To find out whether some sequence you care about, a variant or a promoter fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a prediction as evidence that the factor binds there.

Transcription factor binding sites

Which tracks show transcription factor binding sites?

For human, three tracks cover most needs. All three are in the Regulation group.

Two older tracks still come up in questions. TFBS Conserved shows predicted sites that are conserved across human, mouse and rat, on hg17, hg18 and hg19 only; it has not been updated in many years and there is no hg38 version. ORegAnno is a curated collection of regulatory elements taken from the literature, so it is small but every entry has a citation, on hg19, hg38, mm10 and dm6.

I cannot find my transcription factor in any track. Where else can I look?

First check whether the experiment simply has not been done. ReMap covers the published ChIP-seq experiments that were available when it was built, so if your factor is absent from ReMap there may be no public ChIP-seq for it in that organism. The ReMap website lets you search by target and download the peaks per factor, and that is the quickest way to check.

If the experiment exists but is newer than our tracks, or was done in a cell type we do not carry, you will need to bring the data in yourself. The usual sources are:

How do I display ENCODE data that is not already a track?

You do not need to download anything or write a custom track by hand. The ENCODE portal will open its data in the Genome Browser for you.

  1. Search the ENCODE portal for what you want, narrowing the results with the filters down the left side. Assay title, Target of assay (the factor), Biosample (the cell type or tissue) and Genome assembly are the useful ones.
  2. Click the Visualize button above the result list.
  3. Pick your assembly in the panel that opens, then click UCSC.

The Browser opens with every experiment in your filtered result set loaded as a track hub, so this works just as well for one experiment as for fifty. Restrict the search before you visualize, since a broad filter can attach a very large number of tracks at once.

ENCODE also publishes a hub for each individual experiment, which is handy if you are scripting or want to keep a link in a session. Substitute the accession into this URL:

https://www.encodeproject.org/experiments/ENCSR000AKO/@@hub/hub.txt

and load it from the My Hubs tab of the Track Hubs page, or by appending it to a browser URL as hgTracks?db=hg38&hubUrl= followed by the hub address.

If you would rather place a single file yourself, note that ENCODE distributes peaks as bigBed and signal as bigWig, both of which the Browser reads directly. Copy the file URL from the portal; you do not need to download the file. Then paste one custom track line at Add Custom Tracks:

track type=bigBed name="CTCF K562 peaks" bigDataUrl=https://www.encodeproject.org/files/ENCFF002CEL/@@download/ENCFF002CEL.bigBed

The Browser does not download the whole file. bigBed and bigWig are indexed, so it fetches only the part covering the region you are looking at, which is why a multi-gigabyte signal file opens in a moment. Full instructions are on the custom tracks help page and the track hub help page.

Promoters, enhancers and other elements

Which tracks show promoters?

It depends on what you mean by a promoter, and the tracks disagree enough that it matters.

Which tracks show enhancers and other regulatory elements?

-For the chromatin evidence behind the called elements, see -ENCODE4 Regulation on hg38, which -carries DNase, ATAC-seq, histone -modification and CTCF signal organized by tissue, and the older ENCODE3 Regulation container on +For the chromatin evidence behind the called elements, see ENCODE4 Regulation, +which carries DNase, ATAC-seq, histone modification and CTCF signal organized by tissue, on +hg38 and +mm10. This replaced ENCODE3 +Regulation +as the default in July 2026; ENCODE3 is kept for archival use and is still on hg19 and hg38.

What are cCREs, and which cCRE track should I use?

A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only and so on. The word candidate is doing real work here: these are regions worth testing, not confirmed regulatory elements.

-The ENCODE cCREs container on hg38 holds -several versions. The ENCODE4 -cCREs registry is the current one and should be your default. The ENCODE4 Core -Collection is a smaller, higher-confidence subset. ENCODE3 cCREs is the -earlier release, kept because a great many published analyses used it and coordinates need to -stay reproducible. The container also carries per-biosample subtracks, which is how you restrict -the classification to one cell type. Mouse -mm10 carries the same three versions.

+The ENCODE cCREs container on +hg38 holds several versions. The +ENCODE4 +cCREs registry became the default in July 2026 and is the one to use: 2.3 million human +and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, +higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with +all four core assays. ENCODE3 cCREs is the earlier release, kept for archival +use because a great many published analyses used it and coordinates need to stay reproducible. The +container also carries per-biosample subtracks, which is how you restrict +the classification to one cell type. Mouse +mm10 carries the same three versions.

Working with the data

How do I restrict a search to one cell type or tissue?

-Most of the large regulatory tracks are composites or superTracks holding many subtracks, one -per cell type or experiment, and they arrive with only a summary view turned on. Click the track -name to open its configuration page, where you will find the list of subtracks and, on the -bigger tracks, filters. ReMap, for instance, lets you filter by transcription factor directly in -the track settings, so you can show one factor across all its experiments.

+Most of the large regulatory tracks are collections holding many separate tracks, one per cell +type or experiment, and they arrive with only a summary view turned on. Click the track name to +open its configuration page, where you will find the list and, on the bigger collections, +filters. ReMap, for instance, lets you filter by transcription factor directly in the track +settings, so you can show one factor across all its experiments.

+

+The largest collections use a different configuration page. Where there are hundreds or +thousands of individual experiments, as in the per-experiment tracks of +ENCODE4 Regulation on +hg38, the page shows facets down the +left side and a +paginated table on the right: tick the tissue, assay or biosample you want and the table narrows +to the matching experiments. This is usually a faster way to reach one cell type than reading a +long list.

To find a track rather than configure one, the Track Search page searches track names and descriptions across the whole assembly, which is usually faster than reading through the track groups.

I have a gene. How do I find the factors that regulate it?

No single track answers this, so you have to work outward from the gene.

  1. Navigate to the gene and zoom out far enough to include the surrounding non-coding sequence. Regulatory elements are often tens or hundreds of kilobases away, and the nearest gene to an element is frequently not its target.
  2. Turn on GeneHancer. Its interaction arcs will show which elements have been linked to your gene, including distant ones, which narrows the search from the whole neighborhood to a handful of regions.
  3. Turn on ReMap or TF ChIP and look at which factors have peaks in those regions. This gives you factors that were measured at that position in some cell type.
  4. Check whether any of those cell types are relevant to your biology. A peak in K562 says little about neurons.
  5. If you need candidates in a cell type nobody has assayed, fall back to JASPAR predictions within the GeneHancer elements, and treat the result as hypotheses to test.

To do this systematically, the Data Integrator will intersect two or more tracks and return a table, and the Table Browser will do the same for a region or for a list of genes.

+ +
I have a variant, not a region. What does it do to regulation?
+

+This is a different question from the rest of this page. Everything above annotates regions and +tells you what is where; the tracks here take one base change and tell you what it does. The +same measured and predicted distinction applies, so start with the one track that is +measured.

+

+MPRAVarDB holds 239,028 variants that were put +through a reporter assay and scored for allelic effect, drawn from 18 MPRA studies covering more +than 30 cell lines and more than 30 diseases or traits. The variants come from GWAS and eQTL +fine-mapping, from saturation mutagenesis of 20 disease-associated regulatory elements, and from +smaller focused screens, so the coverage is concentrated on loci people have already had reason +to care about. Items are colored by significance, dark red for FDR below 0.05. If your variant +is in here, you have an experimental answer and not a guess. It lives in the +MPRAs collection, alongside MPRA Base, which tests whole elements. +Available for hg38.

+

+Most variants will not be in it, since 239,028 tested positions is a very small part of the +genome. Two prediction tracks fill the gap by scoring every possible substitution, and neither +is in the Regulation group: look under Phenotype and Disease Associations, in +the Deleteriousness Predictions collection on +hg38.

+ +

+A high score from either says a model thinks the change is disruptive, not that anyone has +measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement +is the better evidence.

+
What is available for assemblies other than human and mouse?

-Considerably less, and this is a matter of what data exists rather than what we have loaded. The -large regulatory resources were built for human first and mouse second. JASPAR predictions are -the most widely available, since they only require the genome sequence and a motif, and are -present for zebrafish, fly, worm, chicken, sea squirt and yeast in addition to human and mouse. -ReMap and ORegAnno cover fly. Most of the remaining tracks described on this page are human and -mouse only.

+Much less. The limit is usually the data itself: these resources were built for human first, and +most never went further than mouse. JASPAR is the main exception, since it needs only the genome +sequence and a motif, so it covers zebrafish, fly, worm, chicken, sea squirt and yeast as well. +ReMap and ORegAnno also cover fly. Everything else described on this page is human and mouse +only.

-Coverage also differs between assemblies of the same organism. On mouse, mm10 carries most of -the regulatory tracks while mm39 has only JASPAR, ReMap and VISTA, because several of the source -projects have not released mm39 versions. If a track you need is on mm10 but not mm39, the -LiftOver tool can convert coordinates between the two, -though you should check the result. Human has the same problem on a smaller scale: a few of -these tracks are still hg19 only, and TFBS Conserved was never rebuilt for hg38.

+Mouse is split awkwardly between its own assemblies. mm10 carries most of these tracks and mm39 +has only JASPAR, ReMap and VISTA, because several of the source projects never released mm39 +versions. If you need something that is on mm10 but not mm39, +LiftOver will convert the coordinates, though check the +result before trusting it. Human has a milder version of the same problem: a few tracks are +still hg19 only, and TFBS Conserved was never rebuilt for hg38.

For assemblies not hosted at UCSC, or for tracks we do not carry, check the public hubs list, where other groups publish data through our browser.

Summary: regulatory tracks by category

Each assembly name below links to that track's description page on that assembly. A few tracks appear on additional genomes not listed here; use Track Search to check a genome that is not shown.

- + - + sacCer3; up to JASPAR 2024 on + hg19 and + mm10 - + - - + + - + - + - + + + + + + + + + + + + + + + + + + + + + +
Track What it is Measured or predicted Assemblies
Transcription factor binding
ReMap ChIP-seq Public ChIP-seq for transcriptional regulators, integrated Measured hg19, hg38, mm10, mm39, dm6
TF ChIP (ENCODE 3 TFBS on hg19) ENCODE 3 TF ChIP-seq peaks, around 340 factors in about 130 cell types Measured hg19, hg38
JASPAR Transcription FactorsMotif matches from the JASPAR CORE collectionMotif matches from the JASPAR CORE collection; several releases as subtracks Predictedhg19, - hg38, - mm10, + JASPAR 2026 on hg38, mm39, danRer11, + galGal6, dm6, ce11, - galGal6, ci3, - sacCer3
TFBS Conserved Conserved motif matches; not updated recently, no hg38 version Predicted hg17, hg18, hg19
ORegAnno Regulatory elements curated from the literature Measured, curated hg19, hg38, mm10, dm6
Promoters and transcription start sites
EPDnew Promoters Experimentally defined promoters with mapped start sites Measured hg19, hg38, mm10
FANTOM5 CAGE transcription start sites and their usage per tissue Measured hg19, hg38, mm10
Enhancers and candidate elements
GeneHancer Regulatory elements linked to predicted target genes Mixed, with predicted targets hg19, hg38
VISTA Enhancers Elements tested individually in transgenic mouse assays Measured, validated hg19, hg38, mm10, mm39
ENCODE cCREsCandidate elements classified from chromatin signalCandidate elements classified from chromatin signal; ENCODE4 is the default Predicted from measured signal hg38, mm10
RefSeq Functional Elements NCBI curated non-coding functional elements Measured, curated hg38, mm10
MPRAsMassively parallel reporter assay activityMPRAs (MPRA Base)Reporter assay activity for 40,938 tested regulatory elements Measured hg38
Chromatin and expression context
ENCODE4 RegulationDNase, ATAC, histone marks and CTCF by tissueDNase, ATAC, histone marks and CTCF by tissue; the current default Measuredhg38hg38, + mm10
ENCODE3 RegulationDNase, histone marks and transcription signalDNase, histone marks and transcription signal; archival, replaced by ENCODE4 Measured hg19, hg38
Single-cell ATAC-seq Accessibility peaks and signal from Cell Browser datasets Measured hg38, mm10
GTEx Gene Gene expression across 53 tissues Measured hg19, hg38
GTEx cis-eQTLs Variants associated with expression of nearby genes Measured hg38
Effect of a single variant
MPRAVarDB239,028 variants tested for allelic effect in reporter assaysMeasuredhg38
AlphaGenomeVariant Impact score for every single-base substitution, coding and non-coding; + under Phenotype and Disease AssociationsPredictedhg38
PromoterAIScore for every single-base substitution in proximal promoters; under Phenotype and + Disease AssociationsPredictedhg38

For the full set of tracks on any assembly, open the Regulation group on the browser page, or use Track Search.