5fd99d465a31d5bd38320ac85ff1969ab9515014 mspeir Wed Sep 23 15:44:06 2026 -0700 Regulation FAQ: correct assembly coverage, add variant effect section, refs #24610 Corrections found by checking the page against our own release announcements, which I should have done before the first commit: - ENCODE4 Regulation is on mm10 as well as hg38. The track symbol differs by assembly (wgEncodeReg4 on hg38, encode4Reg on mm10), so a search by name missed it. - JASPAR 2026 covers hg38, mm39, danRer11, galGal6, dm6, ce11, ci3 and sacCer3. hg19 and mm10 stop at JASPAR 2024. The page had implied all ten were 2026. Also notes that the track holds several releases as subtracks, so the version depends on which one is turned on. - TFBS Conserved is on hg17 as well as hg18 and hg19. - ENCODE4 replaced ENCODE3 as the default in July 2026 and ENCODE3 is kept for archival use. Said in the text and marked in the summary table. - The cell type section described the old subtrack list for the large collections, which now use the faceted interface. - Dropped "composite" and "superTrack", which users should not see. New section on what a single variant does to regulation, covering MPRAVarDB for variants that were actually tested in a reporter assay, and AlphaGenome and PromoterAI for predictions. These live under Phenotype and Disease Associations rather than Regulation, which is worth saying since that is where a reader would otherwise fail to find them. Links to goldenPath/help/hgRegMotifHelp.html from the motif definition and from the JASPAR entry, and the definition now matches the wording there. Co-Authored-By: Claude Opus 5 (1M context) diff --git src/hg/htdocs/FAQ/FAQregulation.html src/hg/htdocs/FAQ/FAQregulation.html index 575f02c1a79..611ee9f2595 100755 --- src/hg/htdocs/FAQ/FAQregulation.html +++ src/hg/htdocs/FAQ/FAQregulation.html @@ -9,30 +9,32 @@

Topics


Return to FAQ Table of Contents

The assembly names after each track are links. They open that track's description page on that assembly, which gives the methods, the data version and the citation. Coverage varies a lot between assemblies, so check the list before you assume a track exists on the genome you work with.

The basics

@@ -41,32 +43,34 @@ them are in the Regulation track group, which you will find below the browser image on the main browser page. It is not always obvious which one you want, because these tracks answer quite different questions even when they all look like boxes on the screen.

What is the difference between measured and predicted binding sites?

Nearly every question we get about these tracks comes down to this distinction.

A measured binding site comes from an experiment, usually ChIP-seq, in which one protein was pulled down in one cell type under one set of conditions. The site is real in the sense that the factor was found there in that experiment. It tells you nothing about other cell types, and the experiment has to have been done for your factor and your tissue for the data to exist at all.

-A predicted binding site comes from scanning the genome sequence for a motif, a short -pattern that the factor is known to prefer. Predictions exist everywhere in the genome for every +A predicted binding site comes from scanning the genome sequence for a +motif, the sequence theme a given factor +prefers, usually stored as a position weight matrix and not a single spelling. Predictions exist +everywhere in the genome for every factor with a known motif, regardless of cell type, and most of them are not bound in vivo. A typical transcription factor motif occurs hundreds of thousands of times in the human genome, while the factor binds only a few thousand of those positions in any given cell.

Neither kind is better than the other. To find out where a factor was actually found, use a measured track. To find out whether some sequence you care about, a variant or a promoter fragment, could plausibly be bound, use a predicted track. What you cannot do is cite a prediction as evidence that the factor binds there.

Transcription factor binding sites

Which tracks show transcription factor binding sites?

For human, three tracks cover most needs. All three are in the Regulation group.

@@ -78,47 +82,51 @@ single project. There are three versions: all peaks per experiment, a non-redundant set that merges similar targets, and cis-regulatory modules. Available for hg19, hg38, mm10, mm39 and dm6.
  • TF ChIP has the ENCODE 3 transcription factor ChIP-seq peaks, all processed the same way. That matters if you are comparing experiments instead of looking up one factor. It covers 340 factors in 129 cell types on hg38, and 338 factors in 130 cell types on hg19, where the track is named ENCODE 3 TFBS rather than TF ChIP. Available for hg19 and hg38. An older clustered version of the - same data is on hg19 as - Txn Factor ChIP.
  • + same data, Txn Factor ChIP, is on + hg19.
  • JASPAR Transcription Factors is the predicted set, made by scanning - the genome with the JASPAR CORE profiles. The current track is JASPAR 2026. Scores are scaled from - 0 to 1000 and reflect how well the sequence matches the profile; the track shows only sites - scoring 400 or above by default, and you can raise that threshold in the track settings. - Available for hg19, + the genome with the JASPAR CORE profiles. Scores are scaled from 0 to 1000 and reflect how well + the sequence matches the profile; the track shows only sites scoring 400 or above by default, + and you can raise that threshold in the track settings. Clicking a site shows the + motif display, with the sequence logo and the + matrix behind the score. The track holds several JASPAR releases as separate subtracks, so check + which one you have turned on before you quote a version. JASPAR 2026 is on hg38, - mm10, mm39, danRer11, + galGal6, dm6, ce11, - galGal6, ci3 and - sacCer3.
  • + sacCer3. On + hg19 and + mm10 the newest release is JASPAR + 2024.

    Two older tracks still come up in questions. TFBS Conserved shows predicted sites that are conserved across human, mouse and rat, on hg17, hg18 and hg19 only; it has not been updated in many years and there is no hg38 version. ORegAnno is a curated collection of regulatory elements taken from the literature, so it is small but every entry has a citation, on hg19, hg38, mm10 and dm6.

    @@ -238,71 +246,87 @@ hg38.
  • VISTA Enhancers has elements that were tested one at a time in transgenic mouse assays, with the resulting expression pattern recorded. It is the smallest of these sets and the best validated. Available for hg19, hg38, mm10 and mm39.
  • RefSeq Functional Elements is NCBI's curated set of experimentally characterized non-coding elements. Available for hg38 and mm10.
  • - MPRAs is a collection of massively parallel reporter assay results. These - measure regulatory activity for many sequences at once. + MPRAs is a collection of massively parallel reporter assay results, which + measure regulatory activity for many sequences at once. The MPRA Base half holds 40,938 tested + elements; the MPRAVarDB half holds tested variants and is covered under + variant effects below. Available for hg38.
  • -For the chromatin evidence behind the called elements, see -ENCODE4 Regulation on hg38, which -carries DNase, ATAC-seq, histone -modification and CTCF signal organized by tissue, and the older ENCODE3 Regulation container on +For the chromatin evidence behind the called elements, see ENCODE4 Regulation, +which carries DNase, ATAC-seq, histone modification and CTCF signal organized by tissue, on +hg38 and +mm10. This replaced ENCODE3 +Regulation +as the default in July 2026; ENCODE3 is kept for archival use and is still on hg19 and hg38.

    What are cCREs, and which cCRE track should I use?

    A candidate cis-regulatory element, or cCRE, is a region that looks regulatory in chromatin data. Nobody has shown that it regulates anything. ENCODE built the Registry of cCREs by combining DNase accessibility with histone modification and CTCF signal across many biosamples, then classifying each region as promoter-like, enhancer-like, CTCF-only and so on. The word candidate is doing real work here: these are regions worth testing, not confirmed regulatory elements.

    -The ENCODE cCREs container on hg38 holds -several versions. The ENCODE4 -cCREs registry is the current one and should be your default. The ENCODE4 Core -Collection is a smaller, higher-confidence subset. ENCODE3 cCREs is the -earlier release, kept because a great many published analyses used it and coordinates need to -stay reproducible. The container also carries per-biosample subtracks, which is how you restrict -the classification to one cell type. Mouse -mm10 carries the same three versions.

    +The ENCODE cCREs container on +hg38 holds several versions. The +ENCODE4 +cCREs registry became the default in July 2026 and is the one to use: 2.3 million human +and 927,000 mouse elements. The ENCODE4 Core Collection is a smaller, +higher-confidence subset, covering the 170 human and 18 mouse biosamples that were profiled with +all four core assays. ENCODE3 cCREs is the earlier release, kept for archival +use because a great many published analyses used it and coordinates need to stay reproducible. The +container also carries per-biosample subtracks, which is how you restrict +the classification to one cell type. Mouse +mm10 carries the same three versions.

    Working with the data

    How do I restrict a search to one cell type or tissue?

    -Most of the large regulatory tracks are composites or superTracks holding many subtracks, one -per cell type or experiment, and they arrive with only a summary view turned on. Click the track -name to open its configuration page, where you will find the list of subtracks and, on the -bigger tracks, filters. ReMap, for instance, lets you filter by transcription factor directly in -the track settings, so you can show one factor across all its experiments.

    +Most of the large regulatory tracks are collections holding many separate tracks, one per cell +type or experiment, and they arrive with only a summary view turned on. Click the track name to +open its configuration page, where you will find the list and, on the bigger collections, +filters. ReMap, for instance, lets you filter by transcription factor directly in the track +settings, so you can show one factor across all its experiments.

    +

    +The largest collections use a different configuration page. Where there are hundreds or +thousands of individual experiments, as in the per-experiment tracks of +ENCODE4 Regulation on +hg38, the page shows facets down the +left side and a +paginated table on the right: tick the tissue, assay or biosample you want and the table narrows +to the matching experiments. This is usually a faster way to reach one cell type than reading a +long list.

    To find a track rather than configure one, the Track Search page searches track names and descriptions across the whole assembly, which is usually faster than reading through the track groups.

    I have a gene. How do I find the factors that regulate it?

    No single track answers this, so you have to work outward from the gene.

    1. Navigate to the gene and zoom out far enough to include the surrounding non-coding sequence. Regulatory elements are often tens or hundreds of kilobases away, and the nearest gene to an element is frequently not its target.
    2. @@ -314,46 +338,89 @@ Turn on ReMap or TF ChIP and look at which factors have peaks in those regions. This gives you factors that were measured at that position in some cell type.
    3. Check whether any of those cell types are relevant to your biology. A peak in K562 says little about neurons.
    4. If you need candidates in a cell type nobody has assayed, fall back to JASPAR predictions within the GeneHancer elements, and treat the result as hypotheses to test.

    To do this systematically, the Data Integrator will intersect two or more tracks and return a table, and the Table Browser will do the same for a region or for a list of genes.

    + +
    I have a variant, not a region. What does it do to regulation?
    +

    +This is a different question from the rest of this page. Everything above annotates regions and +tells you what is where; the tracks here take one base change and tell you what it does. The +same measured and predicted distinction applies, so start with the one track that is +measured.

    +

    +MPRAVarDB holds 239,028 variants that were put +through a reporter assay and scored for allelic effect, drawn from 18 MPRA studies covering more +than 30 cell lines and more than 30 diseases or traits. The variants come from GWAS and eQTL +fine-mapping, from saturation mutagenesis of 20 disease-associated regulatory elements, and from +smaller focused screens, so the coverage is concentrated on loci people have already had reason +to care about. Items are colored by significance, dark red for FDR below 0.05. If your variant +is in here, you have an experimental answer and not a guess. It lives in the +MPRAs collection, alongside MPRA Base, which tests whole elements. +Available for hg38.

    +

    +Most variants will not be in it, since 239,028 tested positions is a very small part of the +genome. Two prediction tracks fill the gap by scoring every possible substitution, and neither +is in the Regulation group: look under Phenotype and Disease Associations, in +the Deleteriousness Predictions collection on +hg38.

    + +

    +A high score from either says a model thinks the change is disruptive, not that anyone has +measured it. Where a variant appears in both MPRAVarDB and a prediction track, the measurement +is the better evidence.

    +
    What is available for assemblies other than human and mouse?

    -Considerably less, and this is a matter of what data exists rather than what we have loaded. The -large regulatory resources were built for human first and mouse second. JASPAR predictions are -the most widely available, since they only require the genome sequence and a motif, and are -present for zebrafish, fly, worm, chicken, sea squirt and yeast in addition to human and mouse. -ReMap and ORegAnno cover fly. Most of the remaining tracks described on this page are human and -mouse only.

    +Much less. The limit is usually the data itself: these resources were built for human first, and +most never went further than mouse. JASPAR is the main exception, since it needs only the genome +sequence and a motif, so it covers zebrafish, fly, worm, chicken, sea squirt and yeast as well. +ReMap and ORegAnno also cover fly. Everything else described on this page is human and mouse +only.

    -Coverage also differs between assemblies of the same organism. On mouse, mm10 carries most of -the regulatory tracks while mm39 has only JASPAR, ReMap and VISTA, because several of the source -projects have not released mm39 versions. If a track you need is on mm10 but not mm39, the -LiftOver tool can convert coordinates between the two, -though you should check the result. Human has the same problem on a smaller scale: a few of -these tracks are still hg19 only, and TFBS Conserved was never rebuilt for hg38.

    +Mouse is split awkwardly between its own assemblies. mm10 carries most of these tracks and mm39 +has only JASPAR, ReMap and VISTA, because several of the source projects never released mm39 +versions. If you need something that is on mm10 but not mm39, +LiftOver will convert the coordinates, though check the +result before trusting it. Human has a milder version of the same problem: a few tracks are +still hg19 only, and TFBS Conserved was never rebuilt for hg38.

    For assemblies not hosted at UCSC, or for tracks we do not carry, check the public hubs list, where other groups publish data through our browser.

    Summary: regulatory tracks by category

    Each assembly name below links to that track's description page on that assembly. A few tracks appear on additional genomes not listed here; use Track Search to check a genome that is not shown.

    @@ -374,42 +441,42 @@ - + - + sacCer3; up to JASPAR 2024 on + hg19 and + mm10 - + - - + + - + - + - + + + + + + + + + + + + + + + + + + + + + +
    hg19, hg38, mm10, mm39, dm6
    TF ChIP (ENCODE 3 TFBS on hg19) ENCODE 3 TF ChIP-seq peaks, around 340 factors in about 130 cell types Measured hg19, hg38
    JASPAR Transcription FactorsMotif matches from the JASPAR CORE collectionMotif matches from the JASPAR CORE collection; several releases as subtracks Predictedhg19, - hg38, - mm10, + JASPAR 2026 on hg38, mm39, danRer11, + galGal6, dm6, ce11, - galGal6, ci3, - sacCer3
    TFBS Conserved Conserved motif matches; not updated recently, no hg38 version Predicted hg17, hg18, hg19
    ORegAnno Regulatory elements curated from the literature Measured, curated hg19, hg38, @@ -444,78 +511,102 @@ Mixed, with predicted targets hg19, hg38
    VISTA Enhancers Elements tested individually in transgenic mouse assays Measured, validated hg19, hg38, mm10, mm39
    ENCODE cCREsCandidate elements classified from chromatin signalCandidate elements classified from chromatin signal; ENCODE4 is the default Predicted from measured signal hg38, mm10
    RefSeq Functional Elements NCBI curated non-coding functional elements Measured, curated hg38, mm10
    MPRAsMassively parallel reporter assay activityMPRAs (MPRA Base)Reporter assay activity for 40,938 tested regulatory elements Measured hg38
    Chromatin and expression context
    ENCODE4 RegulationDNase, ATAC, histone marks and CTCF by tissueDNase, ATAC, histone marks and CTCF by tissue; the current default Measuredhg38hg38, + mm10
    ENCODE3 RegulationDNase, histone marks and transcription signalDNase, histone marks and transcription signal; archival, replaced by ENCODE4 Measured hg19, hg38
    Single-cell ATAC-seq Accessibility peaks and signal from Cell Browser datasets Measured hg38, mm10
    GTEx Gene Gene expression across 53 tissues Measured hg19, hg38
    GTEx cis-eQTLs Variants associated with expression of nearby genes Measured hg38
    Effect of a single variant
    MPRAVarDB239,028 variants tested for allelic effect in reporter assaysMeasuredhg38
    AlphaGenomeVariant Impact score for every single-base substitution, coding and non-coding; + under Phenotype and Disease AssociationsPredictedhg38
    PromoterAIScore for every single-base substitution in proximal promoters; under Phenotype and + Disease AssociationsPredictedhg38

    For the full set of tracks on any assembly, open the Regulation group on the browser page, or use Track Search.