4fa04b06c8c9e1627bb6bedef9e78a02fc4ceb60 mspeir Sat Sep 26 16:25:41 2026 -0700 minor tewaks to wording; removing oreganno refs; removing params from hgTracks links so they use users cart, refs #24610 diff --git src/hg/htdocs/FAQ/FAQregulation.html src/hg/htdocs/FAQ/FAQregulation.html index 611ee9f2595..22999c560f5 100755 --- src/hg/htdocs/FAQ/FAQregulation.html +++ src/hg/htdocs/FAQ/FAQregulation.html @@ -29,38 +29,34 @@
Return to FAQ Table of Contents
The assembly names after each track are links. They open that track's description page on that assembly, which gives the methods, the data version and the citation. Coverage varies a lot between assemblies, so check the list before you assume a track exists on the genome you work with.
The Genome Browser carries a large number of tracks that annotate regulatory regions. Most of them are in the Regulation track group, which you will find below the browser image on -the main browser page. It is not always obvious which -one you want, because these tracks answer quite different questions even when they all look -like boxes on the screen.
+the main browser page.-Nearly every question we get about these tracks comes down to this distinction.
-A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.
A predicted binding site comes from scanning the genome sequence for a motif, the sequence theme a given factor prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.
Neither kind is better than the other. To find out where a factor was actually found, use a @@ -104,55 +100,42 @@ motif display, with the sequence logo and the matrix behind the score. The track holds several JASPAR releases as separate subtracks, so check which one you have turned on before you quote a version. JASPAR 2026 is on hg38, mm39, danRer11, galGal6, dm6, ce11, ci3 and sacCer3. On hg19 and mm10 the newest release is JASPAR 2024. -
-Two older tracks still come up in questions. TFBS Conserved shows predicted -sites that are conserved across human, mouse and rat, on -hg17, -hg18 and -hg19 only; it has not been updated in -many years and -there is no hg38 version. ORegAnno is a curated collection of regulatory -elements taken from the literature, so it is small but every entry has a citation, on -hg19, -hg38, -mm10 and -dm6.
First check whether the experiment simply has not been done. ReMap covers the published ChIP-seq experiments that were available when it was built, so if your factor is absent from ReMap there may be no public ChIP-seq for it in that organism. The ReMap website lets you search by target and download the peaks per factor, and that is the quickest way to check.
If the experiment exists but is newer than our tracks, or was done in a cell type we do not -carry, you will need to bring the data in yourself. The usual sources are:
+carry, you will need to load the data yourself as a custom track. The usual sources are:https://www.encodeproject.org/experiments/ENCSR000AKO/@@hub/hub.txt
and load it from the My Hubs tab of the Track
Hubs page, or by appending it to a browser URL as
hgTracks?db=hg38&hubUrl= followed by the hub address.
If you would rather place a single file yourself, note that ENCODE distributes peaks as bigBed and signal as bigWig, both of which the Browser reads directly. Copy the file URL from the portal; you do not need to download the file. Then paste one custom track line at Add Custom Tracks:
track type=bigBed name="CTCF K562 peaks" bigDataUrl=https://www.encodeproject.org/files/ENCFF002CEL/@@download/ENCFF002CEL.bigBed
-The Browser does not download the whole file. bigBed and bigWig are indexed, so it fetches only -the part covering the region you are looking at, which is why a multi-gigabyte signal file opens -in a moment. Full instructions are on the +Full instructions are on the custom tracks help page and the track hub help page.
It depends on what you mean by a promoter, and the tracks disagree enough that it matters.
A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only -and so on. The word candidate is doing real work here: these are regions worth testing, -not confirmed regulatory elements.
+and so on.The ENCODE cCREs container on hg38 holds several versions. The ENCODE4 cCREs registry became the default in July 2026 and is the one to use: 2.3 million human and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with all four core assays. ENCODE3 cCREs is the earlier release, kept for archival use because a great many published analyses used it and coordinates need to stay reproducible. The container also carries per-biosample subtracks, which is how you restrict the classification to one cell type. Mouse mm10 carries the same three versions.
A high score from either says a model thinks the change is disruptive, not that anyone has measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement is the better evidence.
Much less. The limit is usually the data itself: these resources were built for human first, and most never went further than mouse. JASPAR is the main exception, since it needs only the genome sequence and a motif, so it covers zebrafish, fly, worm, chicken, sea squirt and yeast as well. -ReMap and ORegAnno also cover fly. Everything else described on this page is human and mouse +ReMap also covers fly. Everything else described on this page is human and mouse only.
Mouse is split awkwardly between its own assemblies. mm10 carries most of these tracks and mm39 has only JASPAR, ReMap and VISTA, because several of the source projects never released mm39 versions. If you need something that is on mm10 but not mm39, LiftOver will convert the coordinates, though check the -result before trusting it. Human has a milder version of the same problem: a few tracks are -still hg19 only, and TFBS Conserved was never rebuilt for hg38.
+result before trusting it. Human has a milder version of the same problem: a few tracks, such +as the older clustered Txn Factor ChIP, are still hg19 only.For assemblies not hosted at UCSC, or for tracks we do not carry, check the public hubs list, where other groups publish data through our browser.
Each assembly name below links to that track's description page on that assembly. A few tracks appear on additional genomes not listed here; use Track Search to check a genome that is not shown.
| JASPAR Transcription Factors | Motif matches from the JASPAR CORE collection; several releases as subtracks | Predicted | JASPAR 2026 on hg38, mm39, danRer11, galGal6, dm6, ce11, ci3, sacCer3; up to JASPAR 2024 on hg19 and mm10 | |
| TFBS Conserved | -Conserved motif matches; not updated recently, no hg38 version | -Predicted | -hg17, - hg18, - hg19 | -|
| ORegAnno | -Regulatory elements curated from the literature | -Measured, curated | -hg19, - hg38, - mm10, - dm6 | -|
| Promoters and transcription start sites | ||||
| EPDnew Promoters | Experimentally defined promoters with mapped start sites | Measured | hg19, hg38, mm10 | |
| FANTOM5 | CAGE transcription start sites and their usage per tissue | Measured | @@ -594,19 +557,19 @@Predicted | hg38 |
| PromoterAI | Score for every single-base substitution in proximal promoters; under Phenotype and Disease Associations | Predicted | hg38 | |
For the full set of tracks on any assembly, open the Regulation group on the -browser page, or use +browser page, or use Track Search.