51975d67889e63fbc5fbad0cf1a909ebcca5f25e
max
  Sat Sep 26 14:15:10 2026 -0700
uniprot otto: let a run finish when a few taxa fail, and say where to read about them

A run over hundreds of organisms always has a few that cannot be built: one UniProt
barely annotates, one NCBI has no gene table for, a genome needing more memory than
was reserved. Until now any single one of them aborted the run before the flip, so
24 failures out of 664 taxa kept the other 640 from being published, twice.

--allowFailures=N carries on and publishes when no more than N taxa failed, and
doUpdate.sh passes 10 unless the caller says otherwise, so the monthly cron is no
longer hostage to a handful of awkward organisms. Above the threshold it still
aborts and publishes nothing, which is the right answer when something systemic
has broken.

Nothing is quieter as a result. Every failure is reported with its traceback as
before, and now each one also gets its own file under failedTaxa/<taxId>.log
holding the taxon, its assemblies, the time and the traceback. lastRun.log is
overwritten by the next run and interleaves every taxon, so a failure someone wants
to look at a day later was hard to find in it; these files are not overwritten
except by another failure of the same taxon.

The end-of-run report names each failed taxon with its assemblies and the path to
its log, and prints the --dbs argument to retry exactly those. The failure mail
from doUpdate.sh lists the log paths too.

Checked all three paths: below the threshold the run continues, above it aborts,
with no failures nothing changes. Checked that the cron form gets
--allowFailures=10, that an explicit --allowFailures wins, and that it does not
disturb other arguments.

refs #38300

diff --git src/hg/utils/otto/uniprot/doUpdate.sh src/hg/utils/otto/uniprot/doUpdate.sh
index d2a50a882d5..0db2a0717db 100755
--- src/hg/utils/otto/uniprot/doUpdate.sh
+++ src/hg/utils/otto/uniprot/doUpdate.sh
@@ -1,97 +1,114 @@
 #!/bin/sh
 # configuration setup and cron wrapper for the doUniprot script
 
 cd /hive/data/outside/otto/uniprot || exit 1
 umask 002
 
 #echo WARNING: NOT DOWNLOADING
 #./doUniprot run --skipDownload
 
 runLog=runLog.txt
 
 logRun() {
     echo "`date '+%Y-%m-%d %H:%M:%S'` $*" >> $runLog
 }
 
 # activate the python environment that has the lxml XML parser. Rebuild it with
 # ./makeVenv.sh if this fails.
 if [ ! -f venv/bin/activate ] ; then
     logRun "PREFLIGHT-FAIL no venv"
     echo "UniProt update did not start: venv/bin/activate is missing."
     echo "Rebuild it with: cd /hive/data/outside/otto/uniprot && ./makeVenv.sh"
     exit 1
 fi
 . venv/bin/activate
 
 # Do not spend 35 minutes downloading UniProt only to find out that the parser cannot
 # start. Run it with --help, which imports lxml and then exits, and stop here if that
 # fails. Invoked exactly the way doUniprot invokes it, so this tests the same python.
 if ! ./uniprotToTab --help > /dev/null 2>&1; then
     logRun "PREFLIGHT-FAIL uniprotToTab cannot start"
     echo "UniProt update did not start: ./uniprotToTab cannot be run."
     echo
     echo "The lxml python module does not import. Rebuild the environment with:"
     echo "    cd /hive/data/outside/otto/uniprot && ./makeVenv.sh"
     echo
     ./uniprotToTab --help 2>&1 | tail -20
     exit 1
 fi
 
 # A killed run would otherwise leave a START with no matching line, which reads the same
 # as a run that is still going. Say it was interrupted, and drop the lock file, which
 # doUniprot's own atexit handler does not get to run on a signal.
 trap 'logRun "INTERRUPTED killed by a signal"; rm -f /hive/data/outside/uniProt/current/doUniprot.lock; exit 130' INT TERM HUP
 
 logRun "START"
 # Anything after the first argument is handed to doUniprot, so a hand restart can say
 # "./doUpdate.sh run -p" to skip the download and the multi-day parse and still get the
 # lock handling, the run log and the failure mail. Cron passes only "run".
 # Guard the shift rather than hiding its complaint: shift is a POSIX special built-in, so
 # in a shell that follows that rule (dash, which is /bin/sh on Debian and in most
 # containers) shifting an empty argument list ends the script then and there. The run log
 # would show a START with no END, which reads exactly like a run still in progress.
 if [ $# -gt 0 ]; then
     shift
 fi
+
+# A monthly run covers a hundred or more organisms and there are always a few that cannot be
+# built: an organism UniProt barely annotates, one NCBI has no gene table for, a genome that
+# needs more memory than we reserved. Without this, any one of them stops the other hundred
+# from being published. Failures are still reported in full and each gets its own log under
+# failedTaxa/. Pass --allowFailures yourself to override this default. refs #38300
+case " $* " in
+    *" --allowFailures"*) ;;
+    *) set -- "$@" --allowFailures=10 ;;
+esac
+
+echo "running: ./doUniprot run $*"
 ./doUniprot run "$@" > lastRun.log 2>&1
 exitCode=$?
 trap - INT TERM HUP
 logRun "END exit=$exitCode"
 
 if grep -q "Is a doUniprot process already running" lastRun.log ; then
     # A run from last month, or a hand-started one, is still going, or crashed and left
     # its lock file behind. Say so in one line instead of the failure report below: this
     # is not a broken pipeline, but a stale lock does need someone to look at it.
     logRun "LOCKED another doUniprot run holds the lock file"
     echo "UniProt update skipped: another doUniprot run holds the lock file."
     echo "If nothing is running, remove /hive/data/outside/uniProt/current/doUniprot.lock"
     exit 0
 fi
 
 if [ $exitCode -ne 0 ] ; then
     # lastRun.log is overwritten by the next run, so keep a copy. Without one, a
     # failure that nobody reads leaves no trace on disk at all.
     cp -f lastRun.log lastFail.log
     logRun "FAIL exit=$exitCode log=lastFail.log"
     echo "Big UniProt update FAILED, exit code $exitCode"
     echo
     echo "Full log: /hive/data/outside/otto/uniprot/lastFail.log"
     echo "Restart manually with:"
     echo "    cd /hive/data/outside/otto/uniprot && ./doUniprot run"
     echo "usually with the -p option to skip download and parsing of the gigantic XML."
     echo
     echo "Last 25 lines of the log:"
     tail -25 lastFail.log
+    if [ -d failedTaxa ] ; then
+        echo
+        echo "Per-taxon failure logs, one file each, kept until the next failure of that taxon:"
+        ls -1 failedTaxa/*.log 2>/dev/null | sed "s|^|    /hive/data/outside/otto/uniprot/|"
+    fi
     exit $exitCode
 fi
 
 if grep -q "are not newer than file in" lastRun.log ; then
     # UniProt had no new release this month. This is the normal case for most months,
     # so stay silent: otto crons only mail when something changed or something broke.
     logRun "NOCHANGE no new UniProt release on the server"
     exit 0
 fi
 
 logRun "OK updated to `cat tab/version.txt`"
 echo "Big UniProt update OK"
 echo "Now serving: `cat tab/version.txt`"