d48a1f1935917a45b20ae782f4a0ab53da13e6a9
braney
  Tue Sep 15 12:53:04 2026 -0700
hgTablesTest: skip an oversized page instead of dying inside the allocator, refs #38359

A dense file-backed track can hand back hundreds of megabytes for a single
five-megabyte test region.  hg38 hgdp returned 602MB, which took carefulAlloc
past its 500MB ceiling, and carefulAlloc exits the process where it stands
rather than errAborting, on the grounds that errAbort itself allocates.  So the
run ended with one line on stderr, nothing in the log, and every table still to
come forfeited.  The arm in quickSubmit meant to catch exactly this and name the
track had never once run.

htmlPage now takes an optional ceiling on the response it will read into memory.
Past it the fetch frees what it has read and errAborts naming the url, which the
robot's errCatch turns back into an ordinary return of no page.  The ceiling
defaults to none, which leaves hgNearTest, hgBlatTest and htmlCheck exactly as
they were.  hgTablesTest sets it to 100MB, a fifth of the allocator ceiling: the
dyString roughly doubles as it grows and the old buffer is still live while the
new one fills, and the parsed page then sits alongside its text.

An oversized page is logged and skipped, not counted as an error.  A track that
answers a 5Mb region with 600MB is one this robot cannot test, which is the same
situation the row count screen already catches before submitting; counting it
would put a failure in every weekly run and leave the summary as useless a gate
as the one that never failed.

The log is line buffered now as well.  Finding out that a run died partway
through is what this robot is for, and a block of buffered lines lost on the way
out is part of how the old failure left no trace of which track it was on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

diff --git src/inc/htmlPage.h src/inc/htmlPage.h
index 45e60bbfbe1..015182f07f2 100644
--- src/inc/htmlPage.h
+++ src/inc/htmlPage.h
@@ -200,30 +200,38 @@
 struct slName *htmlPageLinks(struct htmlPage *page);
 /* Scan through tags list and pull out HREF attributes. */
 
 struct slName *htmlPageSrcLinks(struct htmlPage *page);
 /* Scan through tags list and pull out SRC attributes. */
 
 void htmlPageFormOrAbort(struct htmlPage *page);
 /* Aborts if no FORM found */
 
 void htmlPageValidateOrAbort(struct htmlPage *page);
 /* Do some basic validations.  Aborts if there is a problem. */
 
 void htmlPageStrictTagNestCheck(struct htmlPage *page);
 /* Do strict tag nesting check.  Aborts if there is a problem. */
 
+#define HTML_PAGE_TOO_BIG "htmlPage: response larger than the"
+/* Prefix of the errAbort message from a fetch that ran past htmlPageSetMaxSize().
+ * A caller catching that abort recognizes it with startsWith(). */
+
+void htmlPageSetMaxSize(size_t maxSize);
+/* Set a ceiling on the size of a response this module will read into memory.  Past it
+ * the fetch errAborts instead, naming the url.  Zero, the default, means no ceiling. */
+
 char *htmlSlurpWithCookies(char *url, struct htmlCookie *cookies);
 /* Send get message to url with cookies, and return full response as
  * a dyString.  This is not parsed or validated, and includes http
  * header lines.  Typically you'd pass this to htmlPageParse() to
  * get an actual page. */
 
 struct htmlPage *htmlPageParse(char *url, char *fullText);
 /* Parse out page and return.  Warn and return NULL if problem. */
 
 struct htmlPage *htmlPageParseOk(char *url, char *fullText);
 /* Parse out page and return only if status ok. */
 
 struct htmlPage *htmlPageParseNoHead(char *url, char *htmlText);
 /* Parse out page in memory (past http header if any) and return. */