215d7a575ff21f96f5bf8d3d4a3d20f95bd1d95a mspeir Sat Sep 26 19:57:36 2026 -0700 assembly hub docs: fix broken twoBit and BLAT examples, correct mirror IPs, lead with the short hub URL, document hubCheck, refs #37641 diff --git src/hg/htdocs/goldenPath/help/assemblyHubHelp.html src/hg/htdocs/goldenPath/help/assemblyHubHelp.html index 85c811d0ec6..d854232b295 100755 --- src/hg/htdocs/goldenPath/help/assemblyHubHelp.html +++ src/hg/htdocs/goldenPath/help/assemblyHubHelp.html @@ -1,23 +1,23 @@ - + -

Assembly Hub User Guide

+

Assembly Hub User Guide

Overview

An Assembly Data Hub is a set of Internet-accessible data files that define the reference sequence to be used for a browser instance, as well as all the data files that define the annotation for that sequence. Assembly Data Hubs allow researchers to use the UCSC Genome Browser to view their own sequences with associated annotation, without the requirement that UCSC support a browser on that sequence.

Note: if you are working with a genome that has already been submitted to the NCBI Assembly system, it may already be available in the UCSC Genome Browser. Please check the GenArk Assembly Hub collection @@ -25,83 +25,80 @@ UCSC Assembly Request page to request that the genome assembly be added.

Contents

Web Server
Assembly Hub Components
+
Checking Your Hub
Linking to Your Assembly Hub
Building Tracks
Assembly Hub Resources
- -
Adding BLAT Servers
- -

Web Server

To display a novel genome sequence in the UCSC Genome Browser, a web server hosted by the institution (or a free service such as Cyverse) can be used. For environments operating behind a firewall, hub files can also be loaded locally through GBiB to provide access to the UCSC Genome Browser. Hosting hub files over HTTP is strongly recommended, as it is significantly more efficient than FTP. A hierarchical directory structure must then be established to organize the files associated with the genome sequence. For example:

 myHub/ - directory to organize your files on this hub
     hub.txt - primary reference text file to define the hub, refers to:
     genomes.txt - definitions for each genome assembly on this hub
         newOrg1/ - directory of files for this specific genome assembly
             newOrg1.2bit - '2bit' file constructed from your fasta sequence
             description.html - information about this assembly for users
             trackDb.txt - definitions for tracks on this genome assembly
             groups.txt - definitions for track groups on this assembly
             bigWig and bigBed files - data for tracks on this assembly
             external track hub data tracks
 

The hub can be referenced by a URL such as: http://yourLab.yourInstitution.edu/myHub/hub.txt

Assembly Hub Components

- +

hub.txt

The initial file, hub.txt is the primary URL reference for the assembly hub:

Format of the file:

 hub hubName
 shortLabel genome
 longLabel Comment describing this hub contents
 genomesFile genomes.txt
 email contactEmail@institution.edu
 descriptionUrl aboutHub.html
 

@@ -123,31 +120,31 @@

genomes.txt

The genomes.txt file provides references to the genome assemblies and tracks available in the assembly hub.

 genome ricCom1
 trackDb ricCom1/trackDb.txt
 groups ricCom1/groups.txt
 description July 2011 Castor bean
 twoBitPath ricCom1/ricCom1.2bit
 organism Ricinus communis
 defaultPos E09R7372:1000000-2000000
 orderKey 4800
 scientificName Ricinus communis
 htmlPath ricCom1/description.html
 transBlat yourLab.yourInstitution.edu 17777
-blat yourLab.yourInstitution.edu 17777
+blat yourLab.yourInstitution.edu 17779
 isPcr yourLab.yourInstitution.edu 17779
 

Multiple assembly definitions can be included in a single file, separated by blank lines. The file references are relative paths. In this example, the subdirectory ricCom1 contains the files for this specific assembly.

Refer to the Adding Groups to a Track hub section of the Track Hubs help page for more details.

Single-File Track Hub (useOneFile on)

Traditionally, an assembly hub required multiple configuration files (hub.txt, genomes.txt, trackDb.txt, and optionally groups.txt), along with a .2bit file for the sequence. The useOneFile on option simplifies this by consolidating everything into a single configuration file. Note: The single-file format supports one genome assembly per file. For multiple assemblies, use the traditional -multi-file setup.

+multi-file setup. A single-file hub that defines a second genome does not report an error: the +Browser uses only the first genome stanza and attaches every track to it.

Example configuration:

 hub mySingleFileHub
 shortLabel My Single-File Hub
 longLabel An example of a single-file UCSC track hub
 useOneFile on
 email myEmail@example.com
 
 genome hg19
 
 track exampleBigWig
 shortLabel BigWig Coverage
 longLabel Coverage data over hg19
 type bigWig
 visibility full
 bigDataUrl http://myServer.com/data/example.bigWig
 
 track exampleVCF
 shortLabel VCF Variants
 longLabel Variant calls over hg19 region
 type vcfTabix
 visibility pack
 bigDataUrl http://myServer.com/data/example.vcf.gz
 

If your hub requires a reference genome sequence, you can still provide a .2bit file with twoBitPath. Grouping (previously in -groups.txt.) can also be integrated here if needed. +groups.txt). can also be integrated here if needed.

Once hosted on a server, the single configuration file (and associated data files such as .bigWig, .vcf.gz, .2bit) can be loaded into the UCSC Genome -Browser via the My Hubs page.

+Browser via the Connected Hubs tab of the +Track Data Hubs page.

Building Tracks

Tracks are defined in the trackDb.txt file, where each stanza specifies how tracks are displayed (shortLabel, longLabel, color, visibility), along with other information such as the group the track belongs to (referencing groups.txt) and whether additional HTML should be displayed when a user clicks into the track or a track item:

 track gap_
 longLabel Gap
 shortLabel Gap
 priority 11
 visibility dense
 color 0,0,0
 bigDataUrl bbi/ricCom1.gap.bb
@@ -468,69 +472,94 @@
 

Assembly hubs can include a Cytoband track, which allows quicker navigation of chromosomes and displays banding pattern information, if known.

A simple version of the track can be built using the existing chrom.sizes file for your assembly. Banding options include: gneg, gpos25, gpos50, gpos75, gpos100, acen, gvar, or stalk).

Example:

 cat araTha1.chrom.sizes | sort -k1,1 -k2,2n | awk '{print $1,0,$2,$1,"gneg"}' > cytoBandIdeo.bed
 

The resulting BED file can be converted into a BigBed file and associated with an .as definition file (see example) to + target="_blank">example) to inform the browser that this is not a standard BED:

 bedToBigBed -type=bed4 cytoBandIdeo.bed -as=cytoBand.as araTha1.chrom.sizes cytoBandIdeo.bigBed
 

In trackDb.txt, if the track is named cytoBandIdeo (e.g., track cytoBandIdeo), it will automatically load into the assembly hub.

+ +

Checking Your Hub

+

+Before loading a new or edited hub in the Browser, run hubCheck against it. The utility is +available from the downloads +page. It reads the hub the way the Browser does and reports missing files, malformed stanzas, +settings the Browser does not recognize and data files it cannot open:

+
+hubCheck https://yourLab.yourInstitution.edu/myHub/hub.txt
+
+

+Problems are reported with the line number of the stanza they came from, which is usually quicker +than working backwards from a hub that loads but shows nothing. A clean run reports no problems; +purely cosmetic omissions, such as a missing hub description page, come back as warnings rather +than errors.

+

+Note: hubCheck validates the files it is given, so it cannot catch every problem. In +particular it does not flag a single-file hub that defines more than +one genome, which fails silently in the Browser.

+

Linking to Your Assembly Hub

-Direct links to the genome(s) within the assembly hub can then be constructed.

+Once the hub is hosted, you can link straight to a genome inside it. Give the +hubUrl of your hub.txt and the genome you want, and the Browser +loads the hub and opens that assembly.

- +

+For hubs published in GenArk +there is a shorter form still, https://genome.ucsc.edu/h/ followed by the accession, +for example +https://genome.ucsc.edu/h/GCA_030020305.1.

+

+A longer form exists that routes through the hub connect page. It is worth knowing about only when +you need the Browser to rebuild the hub's track list rather than reuse what it has already +cached, which is occasionally useful while a hub is still being developed:

+
+https://genome.ucsc.edu/cgi-bin/hgHubConnect?hgHub_do_redirect=on&hgHubConnect.remakeTrackHub=on&hgHub_do_firstDb=1&hubUrl=yourHubUrl
+

Assembly Hub Resources

Resources for automatically building assembly hubs include G-OnRamp and MakeHub.

G-OnRamp

G-OnRamp is a Galaxy workflow that turns a genome assembly and RNA-Seq data into a Genome Browser with multiple evidence tracks. Since G-OnRamp is based on the Galaxy platform, becoming familiar with Galaxy concepts and functionalities is recommended. See their @@ -670,112 +699,116 @@ isPcr yourServer.yourInstitution.edu 17779

With this configuration, BLAT and PCR searches become available for the assembly. For example:

 http://genome.ucsc.edu/cgi-bin/hgBlat?hubUrl=http://yourServer.yourInstitution.edu/myHub/hub.txt
 

This URL opens the BLAT interface, where the assembly will appear in the Genome drop-down menu. The isPcr line enables the use of a different gfServer instance for PCR queries if desired.

Firewall note: Some institutions block repeated BLAT server queries. In such cases, administrators must whitelist the following IP ranges:

Further details on gfServer options are available from the Source Downloads page (pre-compiled binaries are located in the blat/ directory) and the blat documentation.

gfServers may also be set up within GBiB for local operation; see the GBiB assembly BLAT setup guide for detailed instructions.

To terminate a gfServer instance, run:

-
gfServer stop localhost 17860
+
gfServer stop localhost 17779

Troubleshooting BLAT servers

Errors may occur if translatedBlat and nucleotideBlat port numbers are reversed. A typical message in this case is:

Expecting 6 words from server got 2

If a gfServer instance is started from the same directory as the .2bit file, for example:

 gfServer start localhost 17779 -stepSize=5 contigsRenamed.2bit &

an attempt to run a DNA sequence query through the web-based BLAT tool may return:

 Error in TCP non-blocking connect() 111 - Connection refused
 Operation now in progress
 Sorry, the BLAT/iPCR server seems to be down. Please try again later.
 
  1. Process check
    - Confirm that a gfServer process is running:
  2. -
    ps aux | grep gfServer
    + Confirm that a gfServer process is running: +
    ps aux | grep gfServer
  3. Verify path and filename
    In the genomes.txt, the twoBitPath/filename must match the .2bit file used when starting gfServer. The location of the gfServer instance can be verified by changing into the directory where gfServer was launched and running the appropriate hostname command.
    hostname -i
    This will return an IP address, for example: 132.249.245.79
    Test the connection with telnet: telnet:
    telnet yourIP yourPort
    For example:
    telnet 132.249.245.79 17777
    A successful connection shows:
    Connected to 132.249.245.79
    If Connection refused appears, gfServer may not be running, or the IP/port configuration is incorrect.
    The genomes.txt file should also be checked to confirm that the BLAT line matches the correct IP and port. For example:
    blat 132.249.245.79 17777
    Instead of:
    blat localhost 17777
  4. Check gfServer status
    Request status directly from gfServer:
    gfServer status yourLocation yourPort
    For example:
    gfServer status 132.249.245.79 17777
    - Sample output might look like:
  5. + Sample output might look like:
    -version 36x2
    +version 39x1
    +serverType static
    +version 39x1
     type nucleotide
     host localhost
     port 17777
     tileSize 11
     stepSize 5
     minMatch 2
     pcr requests 0
     blat requests 0
     bases 0
     misses 0
    -noSig 1
    +noSig 0
     trimmed 0
     warnings 0
    -
    +
  6. Test with gfClient
    A reliable troubleshooting method is to bypass the web interface and use the command-line utility gfClient. If gfClient successfully connects to gfServer, the IP/port configuration is correct. Running gfClient directly verifies connectivity independently of the browser interface. From the directory containing the hub's .2bit file, the command can be executed as follows:
    gfClient yourLocation yourPort pathTo2bitFile yourFastaQuery.fa output.psl
    For example:
    gfClient localhost 17777 . query.fa gfOutput.psl
    Note the . after the port, which tells gfClient to use the .2bit file in the current directory. Check gfOutput.psl for BLAT results.