9687a91909838021ee4ca314e2a29c0cccd845c8 mspeir Sat Sep 26 16:32:35 2026 -0700 commenting out the section about loading ENCODE files as CTs, refs #24610 diff --git src/hg/htdocs/FAQ/FAQregulation.html src/hg/htdocs/FAQ/FAQregulation.html index 22999c560f5..debed8a8caa 100755 --- src/hg/htdocs/FAQ/FAQregulation.html +++ src/hg/htdocs/FAQ/FAQregulation.html @@ -1,575 +1,576 @@
Return to FAQ Table of Contents
The assembly names after each track are links. They open that track's description page on that assembly, which gives the methods, the data version and the citation. Coverage varies a lot between assemblies, so check the list before you assume a track exists on the genome you work with.
The Genome Browser carries a large number of tracks that annotate regulatory regions. Most of them are in the Regulation track group, which you will find below the browser image on the main browser page.
A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.
A predicted binding site comes from scanning the genome sequence for a motif, the sequence theme a given factor prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.
Neither kind is better than the other. To find out where a factor was actually found, use a measured track. To find out whether some sequence you care about, a variant or a promoter fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a prediction as evidence that the factor binds there.
For human, three tracks cover most needs. All three are in the Regulation group.
First check whether the experiment simply has not been done. ReMap covers the published ChIP-seq experiments that were available when it was built, so if your factor is absent from ReMap there may be no public ChIP-seq for it in that organism. The ReMap website lets you search by target and download the peaks per factor, and that is the quickest way to check.
If the experiment exists but is newer than our tracks, or was done in a cell type we do not carry, you will need to load the data yourself as a custom track. The usual sources are:
You do not need to download anything or write a custom track by hand. The ENCODE portal will open its data in the Genome Browser for you.
The Browser opens with every experiment in your filtered result set loaded as a track hub, so this works just as well for one experiment as for fifty. Restrict the search before you visualize, since a broad filter can attach a very large number of tracks at once.
ENCODE also publishes a hub for each individual experiment, which is handy if you are scripting or want to keep a link in a session. Substitute the accession into this URL:
https://www.encodeproject.org/experiments/ENCSR000AKO/@@hub/hub.txt
and load it from the My Hubs tab of the Track
Hubs page, or by appending it to a browser URL as
hgTracks?db=hg38&hubUrl= followed by the hub address.
It depends on what you mean by a promoter, and the tracks disagree enough that it matters.
For the chromatin evidence behind the called elements, see ENCODE4 Regulation, which carries DNase, ATAC-seq, histone modification and CTCF signal organized by tissue, on hg38 and mm10. This replaced ENCODE3 Regulation as the default in July 2026; ENCODE3 is kept for archival use and is still on hg19 and hg38.
A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only and so on.
The ENCODE cCREs container on hg38 holds several versions. The ENCODE4 cCREs registry became the default in July 2026 and is the one to use: 2.3 million human and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with all four core assays. ENCODE3 cCREs is the earlier release, kept for archival use because a great many published analyses used it and coordinates need to stay reproducible. The container also carries per-biosample subtracks, which is how you restrict the classification to one cell type. Mouse mm10 carries the same three versions.
Most of the large regulatory tracks are collections holding many separate tracks, one per cell type or experiment, and they arrive with only a summary view turned on. Click the track name to open its configuration page, where you will find the list and, on the bigger collections, filters. ReMap, for instance, lets you filter by transcription factor directly in the track settings, so you can show one factor across all its experiments.
The largest collections use a different configuration page. Where there are hundreds or thousands of individual experiments, as in the per-experiment tracks of ENCODE4 Regulation on hg38, the page shows facets down the left side and a paginated table on the right: tick the tissue, assay or biosample you want and the table narrows to the matching experiments. This is usually a faster way to reach one cell type than reading a long list.
To find a track rather than configure one, the Track Search page searches track names and descriptions across the whole assembly, which is usually faster than reading through the track groups.
No single track answers this, so you have to work outward from the gene.
To do this systematically, the Data Integrator will intersect two or more tracks and return a table, and the Table Browser will do the same for a region or for a list of genes.
This is a different question from the rest of this page. Everything above annotates regions and tells you what is where; the tracks here take one base change and tell you what it does. The same measured and predicted distinction applies, so start with the one track that is measured.
MPRAVarDB holds 239,028 variants that were put through a reporter assay and scored for allelic effect, drawn from 18 MPRA studies covering more than 30 cell lines and more than 30 diseases or traits. The variants come from GWAS and eQTL fine-mapping, from saturation mutagenesis of 20 disease-associated regulatory elements, and from smaller focused screens, so the coverage is concentrated on loci people have already had reason to care about. Items are colored by significance, dark red for FDR below 0.05. If your variant is in here, you have an experimental answer and not a guess. It lives in the MPRAs collection, alongside MPRA Base, which tests whole elements. Available for hg38.
Most variants will not be in it, since 239,028 tested positions is a very small part of the genome. Two prediction tracks fill the gap by scoring every possible substitution, and neither is in the Regulation group: look under Phenotype and Disease Associations, in the Deleteriousness Predictions collection on hg38.
A high score from either says a model thinks the change is disruptive, not that anyone has measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement is the better evidence.
Much less. The limit is usually the data itself: these resources were built for human first, and most never went further than mouse. JASPAR is the main exception, since it needs only the genome sequence and a motif, so it covers zebrafish, fly, worm, chicken, sea squirt and yeast as well. ReMap also covers fly. Everything else described on this page is human and mouse only.
Mouse is split awkwardly between its own assemblies. mm10 carries most of these tracks and mm39 has only JASPAR, ReMap and VISTA, because several of the source projects never released mm39 versions. If you need something that is on mm10 but not mm39, LiftOver will convert the coordinates, though check the result before trusting it. Human has a milder version of the same problem: a few tracks, such as the older clustered Txn Factor ChIP, are still hg19 only.
For assemblies not hosted at UCSC, or for tracks we do not carry, check the public hubs list, where other groups publish data through our browser.
Each assembly name below links to that track's description page on that assembly. A few tracks appear on additional genomes not listed here; use Track Search to check a genome that is not shown.
| Track | What it is | Measured or predicted | Assemblies |
|---|---|---|---|
| Transcription factor binding | |||
| ReMap ChIP-seq | Public ChIP-seq for transcriptional regulators, integrated | Measured | hg19, hg38, mm10, mm39, dm6 |
| TF ChIP (ENCODE 3 TFBS on hg19) | ENCODE 3 TF ChIP-seq peaks, around 340 factors in about 130 cell types | Measured | hg19, hg38 |
| JASPAR Transcription Factors | Motif matches from the JASPAR CORE collection; several releases as subtracks | Predicted | JASPAR 2026 on hg38, mm39, danRer11, galGal6, dm6, ce11, ci3, sacCer3; up to JASPAR 2024 on hg19 and mm10 |
| Promoters and transcription start sites | |||
| EPDnew Promoters | Experimentally defined promoters with mapped start sites | Measured | hg19, hg38, mm10 |
| FANTOM5 | CAGE transcription start sites and their usage per tissue | Measured | hg19, hg38, mm10 |
| Enhancers and candidate elements | |||
| GeneHancer | Regulatory elements linked to predicted target genes | Mixed, with predicted targets | hg19, hg38 |
| VISTA Enhancers | Elements tested individually in transgenic mouse assays | Measured, validated | hg19, hg38, mm10, mm39 |
| ENCODE cCREs | Candidate elements classified from chromatin signal; ENCODE4 is the default | Predicted from measured signal | hg38, mm10 |
| RefSeq Functional Elements | NCBI curated non-coding functional elements | Measured, curated | hg38, mm10 |
| MPRAs (MPRA Base) | Reporter assay activity for 40,938 tested regulatory elements | Measured | hg38 |
| Chromatin and expression context | |||
| ENCODE4 Regulation | DNase, ATAC, histone marks and CTCF by tissue; the current default | Measured | hg38, mm10 |
| ENCODE3 Regulation | DNase, histone marks and transcription signal; archival, replaced by ENCODE4 | Measured | hg19, hg38 |
| Single-cell ATAC-seq | Accessibility peaks and signal from Cell Browser datasets | Measured | hg38, mm10 |
| GTEx Gene | Gene expression across 53 tissues | Measured | hg19, hg38 |
| GTEx cis-eQTLs | Variants associated with expression of nearby genes | Measured | hg38 |
| Effect of a single variant | |||
| MPRAVarDB | 239,028 variants tested for allelic effect in reporter assays | Measured | hg38 |
| AlphaGenome | Variant Impact score for every single-base substitution, coding and non-coding; under Phenotype and Disease Associations | Predicted | hg38 |
| PromoterAI | Score for every single-base substitution in proximal promoters; under Phenotype and Disease Associations | Predicted | hg38 |
For the full set of tracks on any assembly, open the Regulation group on the browser page, or use Track Search.