d1444e2ca8e21460228433810b1a6703c6520db1
braney
  Mon Sep 7 12:02:02 2026 -0700
otto: keep the two surveys beside the monitor that reads them, refs #38101

The monitor's stamp table is the machine-readable half of a survey that
lived only in the redmineNotes directory, which is not a repository. The
reasoning behind each of the forty globs, and the measured curl shape
behind each source URL, were therefore one copy on /hive.

Both surveys now sit beside the script. Read ottoFailureSignatures.tsv
before changing a stamp glob: it says what each job writes and when, which
is the difference between a stamp that tracks every run and one that only
moves when the data changes.

diff --git src/hg/utils/otto/ottoMonitor/ottoFailureSignatures.tsv src/hg/utils/otto/ottoMonitor/ottoFailureSignatures.tsv
new file mode 100644
index 00000000000..5c8ffbdd907
--- /dev/null
+++ src/hg/utils/otto/ottoMonitor/ottoFailureSignatures.tsv
@@ -0,0 +1,78 @@
+# Kept in the kent tree at src/hg/utils/otto/ottoMonitor, beside the monitor
+# that reads it.  A dated copy of the survey it came from is in
+# /hive/groups/browser/redmineNotes/38101/claude/.
+#
+# What a run, a change, and a failure look like on disk, per otto job.
+# For the #38101 failure monitor.  Surveyed 2026-09-07, read-only, from the LIVE
+# copies under /hive/data/outside/otto.  Nothing was edited or run.
+#
+# This covers all 29 jobs that have an external source.  The 18 internal jobs
+# are NOT surveyed yet; see the note beside this file.
+#
+# runStamp: what proves the job RAN, whether or not the data changed.  Without
+# one, "no new data" and "did not run" are the same picture on disk.
+#   log         a log file written on every run
+#   dateDir     a dated directory created before any change check
+#   perRunFile  a named file rewritten every run
+#   dirMtime    NOTHING is written that survives, but a temp file is created and
+#               removed, so the job directory's own mtime tracks the last run.
+#               Works today, but it is a side effect, not a design, and any
+#               manual touch in the directory destroys it.
+#   none        no evidence at all that the job ran
+#
+# changeArtifact: what appears only when the source actually changed.
+# failureSignal: what a failure produces.  Cron mails only on output.
+#
+# job  cadence  runStamp  runStampPath  changeArtifact  failureSignal
+clinvar	monthly 8th 00:08	log	log/clinvar.log, appended by tee -a on every run	bigBed/, prevRun/, dated outputs	set -o errexit -o pipefail; all output goes to the log and to mail
+geneReviews	weekly Tue 08:00	log	lastRun.log	dated dir, ../lastUpdate	ERR trap prints "ERROR: GeneReviews build failed (exit N)"; exit 255 if the work dir is missing
+decipher	daily 04:11	dirMtime	the curl writes and later removes the variants bed	dated dir, release/	ERR trap reportErr
+gwas	weekly Wed 04:41	none	-	dated dir YYMMDD	exit 1 if the column list does not match expectedColumns.txt; exit 255 if the work dir is missing
+dbVar	monthly 2nd 08:13	dateDir	<today>/giab, made before any check	release/<db>/version.txt	ERR trap cleanUpOnError; prints "lastUpdate was NOT bumped" and exits 1
+orphanet	monthly 10th 07:10	dateDir	<today>, made in the wrapper before any check	release/<db>/version.txt	ERR trap; a count change over 10 percent prints and exits 1
+insight	weekly Tue 03:08	log	log/insight.<date>.log	dated build dir, but it is DELETED on a no-change run	set -o errexit -o pipefail; a count change over 10 percent fails
+omim	daily 04:17	dateDir	<today>/wget.log, made before the md5 compare	release.diff, prev.md5sum.txt	prints "Potential error in OMIM release fetch" and cats wget.log
+lovd	weekly Mon 13:14	dateDir	<today>, made at the top of download.sh	release/<db>	exit 255 if the work dir is missing; a bed under 10000 lines is rejected
+mitoMap	weekly Wed 08:55	dirMtime	.latest.tsv files written then removed on no-change	version.txt, dated bb files	downloadFile() exits 1 on a failed or empty fetch; a count change over 20 exits 1
+refSeqHistorical	weekly Wed 08:40	none	-	lastHandledRelease.txt	five retries, then prints "could not reach <url>" and exits 1
+strchive	weekly Mon 07:45	none	-	lastRelease.txt, releases/<tag>/	prints and exits 1 on anything unexpected
+vcepVersions	monthly 15th 06:40	none	-	nothing at all, it is a pure notifier	raises RuntimeError naming the URL and what it could not parse
+uniprot	monthly 26th 07:00	log	lastRun.log, doUniprot.log	the 110+ db outputs	runs 3 to 4 days; a failure mid-run leaves a partial tree
+uniprotWuhCor1	daily 04:00	none	-	archive/covid-19.<date>.xml	NONE. curl -s, exit code ignored. See #38280, source dead since 2023
+clinGen	daily 09:00	perRunFile	clinGenGeneValidity/ls.check and release.list, rewritten every run	release.diff	
+clinGenCspec	weekly Wed 11:11	dirMtime	svis.json written then removed on no-change	svis.json.old, the two bb files	four retries, then a message naming the CSpec status page, exit 1
+varChat	weekly Thu 22:22	dirMtime	.latest.bb files written then removed on no-change	varChat.hg*.bb, version.txt	set -e -o pipefail; a count difference guard
+vista	monthly 1st 09:00	dirMtime	.latest.bed and .latest.bb written then removed on no-change	vista.*.bb	WEAK: no set -e, and bedToBigBed stderr goes to /dev/null
+civic	monthly 2nd 12:12	dirMtime	gene_features/temp* rewritten every run	the civic bb files	set -o pipefail -o errexit
+grcIncidentDb	daily 06:33	perRunFile	the per-assembly directories are touched every run	the incident bb files	
+ncbiRefSeq	weekly Wed 08:23	none	-	prev<Db>.sum	set -beEu -o pipefail; work happens under /hive/data/genomes/<db>/bed
+mane	weekly Mon 05:11	none	-	mane.<version>/	WEAK: the whole body is gated on "if the output bb does not exist"
+malacards	monthly 20th 04:04	perRunFile	geneSymbolToKgId.txt, written from hgsql every run	oldVersions/	
+pubtatorDbSnp	daily 08:28	log	out/<YYYY>/<MM>/<DD>/script_output.log, a full zsh -x trace	rsids.hg*.bed	ERR trap mails the owner. Also mails on NO change, which breaks the silent contract
+panelApp	weekly Tue 10:10	perRunFile	missing_genes.txt, currentJson/	current/<db>	
+g2p	monthly 5th 06:22	perRunFile	doG2p.lock, prevAllG2P.csv	dated dir	
+genCC	weekly Tue 16:08	dirMtime	newSubmission.tsv written then removed	dated dir	the md5 compare happens before anything is written
+trackLists	weekly Thu 06:32	perRunFile	collected.json, mirrorTracks.html, internal.html, all rewritten every run	-	set -o errexit -o pipefail
+#
+# ---- the 18 internal jobs, surveyed 2026-09-07 ----
+# Same columns.  None of these can be source-unreachable, so the only question
+# the monitor asks of them is "did it run".
+#
+readOnlyKentMirror	daily 00:41	perRunFile	~otto/git.fetch.output, overwritten every run	the kent tree reset to origin/master	mails gbauto@ucsc.edu DIRECTLY, not otto-group, then exits 255
+ottoLastLog	month end 23:38	perRunFile	~otto/lastLog/log/<YYYY>/lastLog.<ts>.txt.gz and history.<ts>.txt.gz	same	set -beEu -o pipefail
+ottoGitVsHive	weekdays 01:00	dirMtime	~otto/ottoCrontab.tmp is written then removed at the end of the run	the divergence report on stdout	prints each mismatch, which cron mails to otto-group
+liftRequest	every 7 minutes	perRunFile	ottoRequest.lock, rewritten with the pid every run	a row in the ottoRequest table	prints "hgsql update error" or "sendmail failed" to stderr
+genArkHgcentral	weekly Tue 10:37	perRunFile	genark.tsv, newGenark.tsv, beforeSort.asmList	rows in hgcentraltest.genark	builds a mail file and sends it; each section says SUCCESSFUL or FAILED
+genArkDevList	daily 18:58	perRunFile	pushRR/logs/<YYYY>/<MM>/gbdbGenArk.<host>.<date>.gz	same	set -o errexit; the run takes two to three hours
+genArkPushRR	daily 01:03	perRunFile	pushRR/<host>.todayList.gz and the dated logs beside them	same	a lock file guards against overrunning itself and mails when it trips
+chainTables	daily 03:11	perRunFile	liftOverChain.fromDb.toDb.txt and quickLiftChain.fromDb.toDb.txt	none, it is a checker not a builder	prints both comm diffs to stderr and exits 255
+sessionThumbnails	weekly Thu 01:11	perRunFile	htdocs/thumbNailLinks.html and the fetched images	same	
+tipOfDay	weekdays 05:05	perRunFile	htdocs/tipOfDay.html, and a line appended to tipsDone.txt	same	
+omimUpload	daily 05:15	dirMtime	the upload/ directory mtime	omimTableDump.tgz under htdocs/omimUpload	prints "Process failed" and exits 1. TRAP: upload/*.date mtimes are set by hgsqlTableDate via utime() to the TABLE date, not the run time
+cellsNewsSec	weekdays 00:00	perRunFile	the news json it rewrites	same	not monitored, cells team
+cellsFacets	weekdays 00:00	perRunFile	htdocs-cells-submit/facets/	same	not monitored, cells team
+cellsFacetsJson	weekdays 00:00	perRunFile	htdocs-cells-submit/facets/*.json	same	not monitored, cells team
+cbPingHidden	monthly 1st 08:00	none	-	a draft per dataset in the queue directory	not monitored, cells team; silent when nothing is due
+cbAnnotServer	every minute	none	-	watchdog.log grows only when it restarts something	not monitored; the real health check is whether gunicorn answers on 127.0.0.1:5051
+cellsTusd	@reboot	none	-	-	not monitored; the crontab line is commented out
+cbDeWorker	-	none	-	-	moved 2026-08-22 to otto's crontab on hgcompute-08, not visible from hgwdev