3a5d6583ad4aea08655f9e1408baa712c6339cd1
braney
  Wed Sep 30 09:47:52 2026 -0700
ottoMonitor: compare each otto track on hgwdev with genome.ucsc.edu, refs #38101, #38436

A job can run on time while its output never reaches the RR, because the push
is a separate root cron that the otto run does not see. DECIPHER rebuilt its
track every week from 2022 while decipherAutoPush was commented out, and the
monitor called it on time.

The monitor now reads the "Data last updated at UCSC" line from hgTrackUi on
both hosts. It reports a track whose public copy has stayed behind hgwdev for
more than 8 days. It finds the tracks itself, from the /gbdb links into the
otto area and the trackDb bigDataUrl that names them. ottoMonitorPublic.tsv
lists only the few it cannot find that way, mostly SQL tables. --no-public
skips the check.

diff --git src/hg/utils/otto/ottoMonitor/README src/hg/utils/otto/ottoMonitor/README
index bbc23416771..f941b398daa 100644
--- src/hg/utils/otto/ottoMonitor/README
+++ src/hg/utils/otto/ottoMonitor/README
@@ -1,121 +1,182 @@
 ottoMonitor - a daily check that every otto job is still running.  Refs #38101.
 
 WHY
 
 Otto jobs are silent when their source has published nothing.  That is the
 design, and it is also why a job that stops running is invisible: no output, no
 mail, and the track quietly freezes.  #38280 is the case that prompted this.  A
 daily job whose source URL had disappeared ran about 1,100 times over three
 years without a word, and the wuhCor1 UniProt track has not moved since June
 2023.
 
 WHAT IT ASKS
 
 One question per job: did it run when it was supposed to.  It answers that from
 a run stamp, meaning something the job leaves behind whether or not the data
 changed.  Of the 40 jobs in ottoMonitorStamps.tsv, twenty-four write a log, a
 dated directory, or a named file on every run.  Eight leave only their working
 directory's mtime, because they write a temp file and delete it.  Eight leave
 nothing at all, and those are reported as blind rather than as passing, so the
 gap stays visible instead of reading as good news.
 
 A late job is not automatically somebody's bug.  So a late job with a source URL
 gets that URL fetched.  If the source is down the job is left alone, and only a
 second run in a row that fails the same way becomes a ticket.  If the source
 answers, the failure was something else and it files the same day.
 
+A SECOND QUESTION: DID IT REACH THE PUBLIC SITE
+
+A job can run on time and users can still see old data.  The copy of a job's
+output to the RR is a separate root cron in /etc/crontab, an AutoPush script,
+and the otto run never sees it.  #38436 is the case that prompted this.  The
+DECIPHER job rebuilt its track every week from 2022 to 2026, but the
+decipherAutoPush line had been commented out, so genome.ucsc.edu kept the data
+from 2022-06-19.  The first question called DECIPHER on time the whole while.
+
+WHICH TRACKS
+
+The monitor finds the tracks itself.  It lists every symlink in /gbdb that
+points into /hive/data/outside/otto, gives each one to the watched job whose
+script directory is the longest prefix of the link's target, and looks up the
+hgwdev trackDb track whose bigDataUrl names the link.  A push copies a
+directory, so it keeps one track per job and /gbdb directory: hg38 and hg19
+when the directory has them, otherwise the one database whose file is newest,
+and within that the track whose file is newest.  uniprot alone has links in 140
+databases, which is why it does not check all of them.  A new file or a new
+assembly is picked up without anyone editing a table.
+
+ottoMonitorPublic.tsv lists only what that misses: the SQL tables, which have
+no /gbdb link, grcIncidentDb, whose bigBed is named by a table and not by
+bigDataUrl, and uniprotWuhCor1, whose files sit in a directory no job runs
+from.  New otto tracks are bigBeds, so that list is not expected to grow.
+
+A track that hgwdev has and genome.ucsc.edu does not, such as clinvarCnvAlpha,
+is on hgwdev only on purpose, and -v lists it as not checked.
+
+HOW IT DECIDES
+
+For each track the monitor reads the "Data last updated at UCSC" line from
+hgTrackUi on hgwdev and on genome.ucsc.edu.  When the public date is older, it
+records in state.json when it first saw that, and it reports the track once it
+has stayed behind for more than lagDays, 8 unless the table says otherwise,
+because the pushes for these jobs run weekly or more often.  The clock starts at
+hgwdev's own date on the first sighting.  It restarts whenever the public date
+moves, because a moving date shows the push is working.  That matters for a job
+that rebuilds every day, because its hgwdev copy never matches the public one at
+the time the monitor runs.
+
+This check needs state.json to see a job that rebuilds often.  With --no-state
+the clock starts fresh at hgwdev's date on every run, so a track that hgwdev
+rebuilt this week is never reported however old the public copy is.  DECIPHER
+is like this: it was rebuilt on 2026-09-27, so the cron run first reports it on
+2026-10-05.
+
+The first run, 2026-09-30, found more tracks that are behind.  civic, pubtator
+and strchive have no AutoPush line at all; strchive is waiting for v504, as
+#38268 plans.  hg19 decipher is behind too: hgwdev has 2021-04-04 data and
+genome.ucsc.edu has 2020-12-06, and the crontab comment says hg19 stopped
+being produced.
+
+The /gbdb scan takes about a minute and a half when the directory cache is
+cold.  The page fetches, two per track with a one-second pause, add about a
+minute more.  --no-public skips this check.
+
 THE SURVEYS BEHIND IT
 
   ottoSourceUrls.tsv         where each job's data comes from, the URL to probe,
                              and the curl shape that URL actually answers.  Every
                              one was fetched, not read out of a script.
   ottoFailureSignatures.tsv  what a run, a change and a failure look like on disk
                              for each of the 47 jobs.  ottoMonitorStamps.tsv is
                              the machine-readable part of it; this is the reasoning.
 
 Both surveys cover all 47 jobs in otto.crontab.  The monitor watches 40 of them.
 The other seven are the Cell Browser jobs, which ottoOwners.tsv marks monitor=no
 because the cells team owns them and a GB Bug would be the wrong tracker.
 
 Read the signature survey before changing a stamp glob.  It says what each job
 writes and when, which is the difference between a stamp that tracks every run
 and one that only moves when the data changes.
 
-THE THREE INPUTS
+THE FOUR INPUTS
 
   ottoOwners.tsv         who owns each job, whether to watch it, its source URL,
                          and its schedule.  Canonical copy is in genecats at
                          otto/ottoOwners.tsv.  The team edits it there.
   ottoMonitorStamps.tsv  where each job's run stamp lives.  Kept beside this
                          script, because it is about the monitor and not about
                          ownership.
+  ottoMonitorPublic.tsv  the few tracks the public-site check cannot find in
+                         /gbdb, mostly SQL tables.  Kept beside this script.
   state.json             written under /hive/data/outside/otto/ottoMonitor/.
                          Holds the last run seen, the last verdict, the number of
                          runs in a row that failed the same way, and the ticket
                          number when one is open.  Without it a second strike
-                         cannot be told from a first.
+                         cannot be told from a first.  The public-site check
+                         keeps its own entries under the key _publicSite.
 
 RUNNING IT
 
   ./ottoMonitor.py                     silent unless something is late
   ./ottoMonitor.py -v                  also lists the on-time jobs and the blind ones
   ./ottoMonitor.py --job clinvar       one job
+  ./ottoMonitor.py --no-public         skip the hgwdev-versus-public comparison
   ./ottoMonitor.py --file              actually file tickets.  OFF by default
 
 Filing is off by default on purpose.  Run it without --file for a while and read
 what it would have filed.
 
 THE TWO COPIES
 
 Same rule as the rest of otto.  Edit the copy in the kent tree at
 src/hg/utils/otto/ottoMonitor, commit, then copy it out to
 /hive/data/outside/otto/ottoMonitor, which is what cron runs.  Never edit the
 hive copy directly.
 
 TRAPS WORTH KNOWING BEFORE YOU CHANGE A STAMP
 
 A file's mtime is not always the run time.  hgsqlTableDate calls utime() to set
 a file's mtime to the database TABLE's date, on purpose, so a later -nt test asks
 about the data rather than the run.  omimUpload's upload/*.date files look like
 an ideal run stamp and are not; the directory mtime is used instead.
 
 A run stamp says the job started, not that it worked.  Most of these globs are
 written early in a run on purpose, so a job that dies half way still shows that
 it tried.  The cost is that a job which fails the same way every time looks on
 time forever.  uniprot is the live example: its monthly run dies after about 37
 minutes on a missing lxml, it has produced no output since January 2025, and the
 browser serves UniProt release 2024_06 while the downloaded source is at 2026_02.
 The monitor still reads it as on time, because lastRun.log is fresh every month.
 Catching that needs a second kind of check, on the age of the job's OUTPUT
 rather than on its schedule, and the monitor does not have one.
 
 The eight directory-mtime jobs work today but nobody designed that signal.
 Copying a file into one of those directories by hand resets the clock and the
 monitor will read it as a run.
 
 The grace window is measured from the last scheduled time whose window has
 already closed, not from the latest scheduled time.  Written the other way, a
 daily job scheduled fewer than graceHours before the monitor's own run time can
 never be called late, because every check lands inside a fresh window.  Six of
 the forty were in that hole: clinGen, genArkPushRR, grcIncidentDb, liftRequest,
 omim and pubtatorDbSnp.
 
 cron's day rule is an OR, not an AND: when both day-of-month and day-of-week are
 restricted the job runs when EITHER matches.  A job like "14 13 * * mon" would
 look like it never runs if this were read the other way.
 
 A job with more than one crontab line has its schedules joined with ";" in the
 cron column of ottoOwners.tsv, and the monitor takes whichever fired last.
 ottoLastLog is the only one so far, because cron cannot say "the last day of the
 month" in one line.
 
 THE EIGHT BLIND JOBS
 
 gwas, mane, ncbiRefSeq, ottoGitVsHive, refSeqHistorical, strchive,
 uniprotWuhCor1, vcepVersions.
 
 Two of them are a one-line fix rather than a monitor problem.  gwas already
 creates a dated directory, just after its change check instead of before.  mane
 gates its whole body on whether an output file already exists.  Moving either
 would put that job in reach.  That is a suggestion for the owners, not something
 this ticket does.