a8900aeeb2ac10fba9a4c772f270d15c54279b87 mspeir Sat Sep 26 18:55:12 2026 -0700 rewrite the example assembly hub sections around GenArk instead of genome-test, refs #37641 diff --git src/hg/htdocs/goldenPath/help/assemblyHubHelp.html src/hg/htdocs/goldenPath/help/assemblyHubHelp.html index 39f1d17cb4c..85c811d0ec6 100755 --- src/hg/htdocs/goldenPath/help/assemblyHubHelp.html +++ src/hg/htdocs/goldenPath/help/assemblyHubHelp.html @@ -35,30 +35,31 @@
  • 2bit File
  • chromAlias
  • groups.txt
  • Single-File Track Hub
  • Linking to Your Assembly Hub
    Building Tracks
    Assembly Hub Resources
    Adding BLAT Servers

    Web Server

    @@ -535,270 +536,118 @@ with Galaxy concepts and functionalities is recommended. See their instruction page for an overview.

    MakeHub

    MakeHub is a command-line tool for fully automatic generation of track data hubs for visualizing genomes with the UCSC Genome Browser. More information is available on their GitHub page.

    Example NCBI assembly hubs

    -There is a collection of example NCBI assembly hubs that can be used directly or copied as -templates. A large collection of script-generated assembly hubs can be browsed on the development server, with -links defaulting to the genome-test site. To load these hubs on the public UCSC site, copy -the hub.txt link and replace the test server domain with the public domain.

    -

    -The following table provides links to launch various assembly hubs grouped by species subsets. By -scrolling down each page, you can access rows for individual assemblies (or groups of assemblies, -e.g., bacteria). Clicking the "common name" hyperlink (e.g., "African bush -elephant" on the Vertebrate Mammalian page) loads the selected hub.

    -
    - - -

    These assemblies use NCBI accession naming patterns. Prototype gene tracks from NCBI gene -predictions are available for a few assemblies. No BLAT servers are provided. Users can copy the -skeleton structure of a hub to run their own BLAT server locally. Brief instructions are available -on each assembly gateway page under "Download files for this assembly hub." +UCSC builds assembly hubs from NCBI GenBank and RefSeq genomes and publishes them as +GenArk, the UCSC Genome +Repository. Every hub there is public. You can load one in the Genome Browser as it stands, or +copy its structure as a working template for a hub of your own.

    +

    +The GenArk index groups assemblies by clade, covering primates, mammals, birds, fishes, other +vertebrates, invertebrates, plants, fungi, viruses, archaea and bacteria. It also lists curated +collections drawn from those clades, among them the Vertebrate Genomes Project, the California +Conservation Genomics Project and the Human Pangenome Reference Consortium. Assemblies that a +newer version has superseded move to a separate legacy collection. New assemblies are added +continuously, so the index is the place to check what exists today, not any list copied onto +this page.

    +

    +Each clade page gives one row per assembly, with a common name, a scientific name, an NCBI +accession, a BioSample, a BioProject and a release date. The common name links to the assembly +open in the Genome Browser, and the scientific name links to the directory of files that make up +the hub.

    +

    +These hubs follow NCBI accession naming patterns, so the genome name is an accession such +as GCA_030020305.1 instead of a UCSC database name such as hg38. Gene +predictions from NCBI RefSeq, Augustus and other sources are available for many of the +assemblies. BLAT and In-Silico PCR run on a shared dynamic gfServer, which is how +tens of thousands of assemblies can offer BLAT without a dedicated server each. If you want the +same arrangement for your own hub, see +Configuring assembly hubs to use a dynamic gfServer +below.

    -

    Example: Loading the African bush elephant assembly hub and reviewing the related genomes.txt - and trackDb.txt

    +

    Example: loading an assembly hub and reading the hub.txt behind it

    -Here are some quick steps to load an example hub from this collection, along with an explanation -of how to view the files behind the hub.

    +These steps load one GenArk assembly, the African savanna elephant, and then open the files it +was built from. Every other assembly in the index works the same way.

      -
    1. Click the - Vertebrate Mammalian assembly hub link above.
    2. -
    3. Scroll down to the common name column and click the hyperlink for - "African bush elephant".
    4. -
    5. You will arrive at a gateway page titled "African bush elephant Genome Browser - - GCA_000001905.1_Loxafr3.0 assembly". This page includes a section, - Data file downloads, where you can access the underlying - files.
    6. -
    7. Click Go (or use the top Genome Browser blue bar menu) to view this assembly hub. - (Note: this will open on our genome-test site.).
    8. -
    9. To load this hub on our public site, copy the hyperlink for - African bush elephant and paste it into your browser. - Then, change the beginning of the URL from
    10. -
      -https://genome-test.gi.ucsc.edu/...
      -
      - to -
      -https://genome.ucsc.edu/...
      -
      +
    11. Open the GenArk index + and click mammals.
    12. +
    13. Search the page for African savanna elephant (hap1 mLoxAfr1 2023), accession + GCA_030020305.1.
    14. +
    15. Click the common name to open the assembly in the Genome Browser. The same assembly is + also reachable at + https://genome.ucsc.edu/h/GCA_030020305.1, a short form of the + hub URL that works for any GenArk accession.
    16. +
    17. Back on the index, click the scientific name to open the + directory of files behind the hub. The path splits the + accession three digits at a time, so GCA_030020305.1 lives under + hubs/GCA/030/020/305/.
    + +

    Exploring the files behind the hub

    -To better understand how the hub works, you can review the associated files:

    -
      -
    1. Go to the GCA_000001905.1_Loxafr3.0 directory - link.
    2. -
    3. Locate the file GCA_000001905.1_Loxafr3.0.ncbi.2bit. This binary indexed file allows - the Browser to display the genome sequence.
    4. -
    5. Open GCA_000001905.1_Loxafr3.0.genomes.ncbi.txt. This genomes.txt file - defines each assembly in the hub. It points to the genome's .2bit file - (twoBitPath) and specifies the trackDb file that contains the - track definitions. (In the case of this large hub with 204 assemblies, the main - genomes.txt file is one directory up, and this stanza is included there.)
    6. -
    7. Review GCA_000001905.1_Loxafr3.0.trackDb.ncbi.txt. This trackDb.txt - file defines the tracks displayed in the hub. It contains bigDataUrl lines - that tell the Browser where to retrieve data for each track, along with optional - settings such as:
    8. +That directory holds the complete hub. Reading a few of its files shows the components described +earlier on this page at work in a real example:

      -
    +

    +A trackDb.txt holding the same track definitions sits next to hub.txt in the +directory. The assemblies also ship files that are not part of the hub definition, such as +GCA_030020305.1.fa.gz, the AGP, and RepeatMasker and RepeatModeler output.

    Adding BLAT servers

    BLAT servers (gfServer) can be configured as either dedicated or dynamic: