a8900aeeb2ac10fba9a4c772f270d15c54279b87 mspeir Sat Sep 26 18:55:12 2026 -0700 rewrite the example assembly hub sections around GenArk instead of genome-test, refs #37641 diff --git src/hg/htdocs/goldenPath/help/assemblyHubHelp.html src/hg/htdocs/goldenPath/help/assemblyHubHelp.html index 39f1d17cb4c..85c811d0ec6 100755 --- src/hg/htdocs/goldenPath/help/assemblyHubHelp.html +++ src/hg/htdocs/goldenPath/help/assemblyHubHelp.html @@ -35,30 +35,31 @@
MakeHub is a command-line tool for fully automatic generation of track data hubs for visualizing genomes with the UCSC Genome Browser. More information is available on their GitHub page.
-There is a collection of example NCBI assembly hubs that can be used directly or copied as -templates. A large collection of script-generated assembly hubs can be browsed on the development server, with -links defaulting to the genome-test site. To load these hubs on the public UCSC site, copy -the hub.txt link and replace the test server domain with the public domain.
--The following table provides links to launch various assembly hubs grouped by species subsets. By -scrolling down each page, you can access rows for individual assemblies (or groups of assemblies, -e.g., bacteria). Clicking the "common name" hyperlink (e.g., "African bush -elephant" on the Vertebrate Mammalian page) loads the selected hub.
- - - -These assemblies use NCBI accession naming patterns. Prototype gene tracks from NCBI gene -predictions are available for a few assemblies. No BLAT servers are provided. Users can copy the -skeleton structure of a hub to run their own BLAT server locally. Brief instructions are available -on each assembly gateway page under "Download files for this assembly hub." +UCSC builds assembly hubs from NCBI GenBank and RefSeq genomes and publishes them as +GenArk, the UCSC Genome +Repository. Every hub there is public. You can load one in the Genome Browser as it stands, or +copy its structure as a working template for a hub of your own.
++The GenArk index groups assemblies by clade, covering primates, mammals, birds, fishes, other +vertebrates, invertebrates, plants, fungi, viruses, archaea and bacteria. It also lists curated +collections drawn from those clades, among them the Vertebrate Genomes Project, the California +Conservation Genomics Project and the Human Pangenome Reference Consortium. Assemblies that a +newer version has superseded move to a separate legacy collection. New assemblies are added +continuously, so the index is the place to check what exists today, not any list copied onto +this page.
++Each clade page gives one row per assembly, with a common name, a scientific name, an NCBI +accession, a BioSample, a BioProject and a release date. The common name links to the assembly +open in the Genome Browser, and the scientific name links to the directory of files that make up +the hub.
+
+These hubs follow NCBI accession naming patterns, so the genome name is an accession such
+as GCA_030020305.1 instead of a UCSC database name such as hg38. Gene
+predictions from NCBI RefSeq, Augustus and other sources are available for many of the
+assemblies. BLAT and In-Silico PCR run on a shared dynamic gfServer, which is how
+tens of thousands of assemblies can offer BLAT without a dedicated server each. If you want the
+same arrangement for your own hub, see
+Configuring assembly hubs to use a dynamic gfServer
+below.
-Here are some quick steps to load an example hub from this collection, along with an explanation -of how to view the files behind the hub.
+These steps load one GenArk assembly, the African savanna elephant, and then open the files it +was built from. Every other assembly in the index works the same way.-https://genome-test.gi.ucsc.edu/... -- to -
-https://genome.ucsc.edu/... -+
GCA_030020305.1.GCA_030020305.1 lives under
+ hubs/GCA/030/020/305/.-To better understand how the hub works, you can review the associated files:
-genomes.txt file
- defines each assembly in the hub. It points to the genome's .2bit file
- (twoBitPath) and specifies the trackDb file that contains the
- track definitions. (In the case of this large hub with 204 assemblies, the main
- genomes.txt file is one directory up, and this stanza is included there.)trackDb.txt
- file defines the tracks displayed in the hub. It contains bigDataUrl lines
- that tell the Browser where to retrieve data for each track, along with optional
- settings such as:useOneFile on, so the hub stanza, the genome stanza and every track
+ stanza sit in this one file instead of being split across hub.txt,
+ genomes.txt and trackDb.txt.twoBitPath points at
+ GCA_030020305.1.2bit, chromSizes and
+ chromAliasBb supply chromosome sizes and
+ alias names, and defaultPos sets the position
+ the browser opens on.blat, transBlat and isPcr lines in that stanza
+ name a dynamic gfServer together with the assembly's path, which is what
+ lets one server answer for many assemblies.bigDataUrl naming a file under
+ bbi/, the directory holding the bigBed and bigWig files with the actual
+ data.ixIxx/;
+ url and
urlLabel: create outbound links to external
- resourceshtml/ to a track.
+
+A trackDb.txt holding the same track definitions sits next to hub.txt in the
+directory. The assemblies also ship files that are not part of the hub definition, such as
+GCA_030020305.1.fa.gz, the AGP, and RepeatMasker and RepeatModeler output.
BLAT servers (gfServer) can be configured as either dedicated or
dynamic: