b3aef6de6abdf2356bf7d4a64fd5d857726eea6b mspeir Sun Oct 4 09:57:49 2026 -0700 Regulation FAQ: cis-regulation background and six interpretive questions Max asked for a short introduction to cis-regulation, and Lou noted that the page leaned towards listing datasets rather than explaining how to read them. This covers both. A new opening question walks through the kinds of evidence the Browser carries for regulation: open chromatin, histone marks, DNA methylation, transcription factor binding, conservation and physical contact. It is adapted from Max's draft, trimmed where it repeated the sections below it. The top of the page now says when the track list was last checked, and CADD 1.7 goes ahead of AlphaGenome among the variant scores, since it has been stable for years. Six questions are new, each one chosen because people keep asking it on the genome list: why a motif turns up across a whole gene, what a GeneHancer arc does and does not claim, what the scores and grey shading mean (which differs from track to track), whether signal heights can be compared (ENCODE4 auto-scales and ENCODE3 does not, which nothing else documents), what an empty region means, and what to do when your tissue was never assayed. The three "Which tracks show ..." headings become "How do I find ...", which is the form the other FAQ pages use. The Single-cell ATAC-seq row leaves the summary table because singleCellSignalsPeaks is still release alpha and those links error anywhere but hgwdev; the row is kept in a comment to restore when the track is released. refs #24610 Co-Authored-By: Claude Opus 5 (1M context) diff --git src/hg/htdocs/FAQ/FAQregulation.html src/hg/htdocs/FAQ/FAQregulation.html index 8e06956e0c6..64febc9776f 100755 --- src/hg/htdocs/FAQ/FAQregulation.html +++ src/hg/htdocs/FAQ/FAQregulation.html @@ -1,85 +1,169 @@

Frequently Asked Questions: Regulation and cis-regulatory tracks

Topics


Return to FAQ Table of Contents

The assembly names after each track are links. They open that track's description page on -that assembly, which gives the methods, the data version and the citation. Coverage varies a +that assembly, which gives the methods, the data version, and the citation. Coverage varies a lot between assemblies, so check the list before you assume a track exists on the genome you work with.

+

+The tracks described here were last checked in October 2026. Projects such as JASPAR and +ENCODE issue new versions on their own schedules, and we add tracks between checks, so use +Track Search if you need to know what +is on an assembly today.

- +

The basics

The Genome Browser carries a large number of tracks that annotate regulatory regions. Most of them are in the Regulation track group, which you will find below the browser image on the main browser page.

+
How do I tell whether a region is regulatory?
+

+This page is about regulation at the level of DNA and chromatin. Mechanisms acting on RNA after +it has been transcribed are a separate subject and are not covered. None of the tracks below +observes regulation directly. Each one measures a property that regulatory regions tend to have, +and the case for any particular region is built by stacking several of them. Almost all of these +signals differ between cell types, so pick the tissue that matters for your question instead of +reading a genome-wide summary.

+ + +
What is the difference between measured and predicted binding sites?

A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.

A predicted binding site comes from scanning the genome sequence for a motif, the sequence theme a given factor prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.

Neither kind is better than the other. To find out where a factor was actually found, use a measured track. To find out whether some sequence you care about, a variant or a promoter fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a prediction as evidence that the factor binds there.

Transcription factor binding sites

-
Which tracks show transcription factor binding sites?
+
How do I find transcription factor binding sites?

For human, three tracks cover most needs. All three are in the Regulation group.

+ +
The JASPAR track shows my factor binding across the whole gene. Why?
+

+This is the expected result for a motif search, not a sign that anything has gone wrong. Motifs +are short and they tolerate mismatches, so a typical one occurs hundreds of thousands of times +in the human genome, while the factor binds only a few thousand of those places in any given +cell.

+

+Turning a matrix into a list of hits means picking a threshold, and there is no standard one. +Two tools working from the same matrix will hand you different sites, because they cut at +different places and weigh conservation differently. This has been an open problem for as long +as people have been scanning genomes, so any particular set of predicted sites is best treated as +one possible answer.

+

+Raising the score cutoff in the track settings will thin the display, but it does not make the +survivors bound. A more productive approach is to read the motif as the list of places the +factor could bind, then narrow that list using evidence that it did bind: a ChIP peak +in a cell type close to yours, open chromatin, or both.

+
I cannot find my transcription factor in any track. Where else can I look?

First check whether the experiment simply has not been done. ReMap covers the published ChIP-seq experiments that were available when it was built, so if your factor is absent from ReMap there may be no public ChIP-seq for it in that organism. The ReMap website lets you search by target and download the peaks per factor, and that is the quickest way to check.

If the experiment exists but is newer than our tracks, or was done in a cell type we do not carry, you will need to load the data yourself as a custom track. The usual sources are: