36668c67d3b1b7ef9da683a2473ab1dea6be6709
braney
  Mon Sep 7 11:52:39 2026 -0700
otto: a daily check that every otto job is still running, refs #38101

Otto jobs are silent when the source has published nothing, which is the
design and also why a job that stops running is invisible. #38280 is the
case that prompted this: a daily job whose source URL had disappeared ran
about 1,100 times over three years without a word.

The monitor asks one question per job, did it run when it was supposed to,
and answers it from a run stamp, meaning something the job leaves behind
whether or not the data changed. Where a job leaves nothing it is reported
as blind rather than as passing, so the gap stays visible.

A late job is not automatically a bug, so a late job with a source URL gets
that URL fetched. A dead source has to fail twice in a row before it becomes
a ticket. A live source means the failure was something else, and that files
the same day.

Filing is off unless --file is given, and the script is silent when every
job is on time.

ottoMonitorStamps.tsv carries the per-job run stamp. It comes from the
survey in /hive/groups/browser/redmineNotes/38101/claude/.

diff --git src/hg/utils/otto/ottoMonitor/README src/hg/utils/otto/ottoMonitor/README
new file mode 100644
index 00000000000..1598ff33e54
--- /dev/null
+++ src/hg/utils/otto/ottoMonitor/README
@@ -0,0 +1,86 @@
+ottoMonitor - a daily check that every otto job is still running.  Refs #38101.
+
+WHY
+
+Otto jobs are silent when their source has published nothing.  That is the
+design, and it is also why a job that stops running is invisible: no output, no
+mail, and the track quietly freezes.  #38280 is the case that prompted this.  A
+daily job whose source URL had disappeared ran about 1,100 times over three
+years without a word, and the wuhCor1 UniProt track has not moved since June
+2023.
+
+WHAT IT ASKS
+
+One question per job: did it run when it was supposed to.  It answers that from
+a run stamp, meaning something the job leaves behind whether or not the data
+changed.  Fifteen jobs write a log or a named file on every run.  Nine leave only
+their working directory's mtime, because they write a temp file and delete it.
+Eight leave nothing at all, and those are reported as blind rather than as
+passing, so the gap stays visible instead of reading as good news.
+
+A late job is not automatically somebody's bug.  So a late job with a source URL
+gets that URL fetched.  If the source is down the job is left alone, and only a
+second run in a row that fails the same way becomes a ticket.  If the source
+answers, the failure was something else and it files the same day.
+
+THE THREE INPUTS
+
+  ottoOwners.tsv         who owns each job, whether to watch it, its source URL,
+                         and its schedule.  Canonical copy is in genecats at
+                         otto/ottoOwners.tsv.  The team edits it there.
+  ottoMonitorStamps.tsv  where each job's run stamp lives.  Kept beside this
+                         script, because it is about the monitor and not about
+                         ownership.
+  state.json             written under /hive/data/outside/otto/ottoMonitor/.
+                         Holds the last run seen, the last verdict, the number of
+                         runs in a row that failed the same way, and the ticket
+                         number when one is open.  Without it a second strike
+                         cannot be told from a first.
+
+RUNNING IT
+
+  ./ottoMonitor.py                     silent unless something is late
+  ./ottoMonitor.py -v                  also lists the on-time jobs and the blind ones
+  ./ottoMonitor.py --job clinvar       one job
+  ./ottoMonitor.py --file              actually file tickets.  OFF by default
+
+Filing is off by default on purpose.  Run it without --file for a while and read
+what it would have filed.
+
+THE TWO COPIES
+
+Same rule as the rest of otto.  Edit the copy in the kent tree at
+src/hg/utils/otto/ottoMonitor, commit, then copy it out to
+/hive/data/outside/otto/ottoMonitor, which is what cron runs.  Never edit the
+hive copy directly.
+
+TRAPS WORTH KNOWING BEFORE YOU CHANGE A STAMP
+
+A file's mtime is not always the run time.  hgsqlTableDate calls utime() to set
+a file's mtime to the database TABLE's date, on purpose, so a later -nt test asks
+about the data rather than the run.  omimUpload's upload/*.date files look like
+an ideal run stamp and are not; the directory mtime is used instead.
+
+The nine directory-mtime jobs work today but nobody designed that signal.
+Copying a file into one of those directories by hand resets the clock and the
+monitor will read it as a run.
+
+cron's day rule is an OR, not an AND: when both day-of-month and day-of-week are
+restricted the job runs when EITHER matches.  A job like "14 13 * * mon" would
+look like it never runs if this were read the other way.
+
+A job with more than one crontab line has its schedules joined with ";" in the
+cron column of ottoOwners.tsv, and the monitor takes whichever fired last.
+ottoLastLog is the only one so far, because cron cannot say "the last day of the
+month" in one line.
+
+THE EIGHT BLIND JOBS
+
+gwas, mane, ncbiRefSeq, ottoGitVsHive, refSeqHistorical, strchive,
+uniprotWuhCor1, vcepVersions.
+
+Two of them are a one-line fix rather than a monitor problem.  gwas already
+creates a dated directory, just after its change check instead of before.  mane
+gates its whole body on whether an output file already exists.  Moving either
+would put that job in reach.  That is a suggestion for the owners, not something
+this ticket does.