d48a1f1935917a45b20ae782f4a0ab53da13e6a9
braney
  Tue Sep 15 12:53:04 2026 -0700
hgTablesTest: skip an oversized page instead of dying inside the allocator, refs #38359

A dense file-backed track can hand back hundreds of megabytes for a single
five-megabyte test region.  hg38 hgdp returned 602MB, which took carefulAlloc
past its 500MB ceiling, and carefulAlloc exits the process where it stands
rather than errAborting, on the grounds that errAbort itself allocates.  So the
run ended with one line on stderr, nothing in the log, and every table still to
come forfeited.  The arm in quickSubmit meant to catch exactly this and name the
track had never once run.

htmlPage now takes an optional ceiling on the response it will read into memory.
Past it the fetch frees what it has read and errAborts naming the url, which the
robot's errCatch turns back into an ordinary return of no page.  The ceiling
defaults to none, which leaves hgNearTest, hgBlatTest and htmlCheck exactly as
they were.  hgTablesTest sets it to 100MB, a fifth of the allocator ceiling: the
dyString roughly doubles as it grows and the old buffer is still live while the
new one fills, and the parsed page then sits alongside its text.

An oversized page is logged and skipped, not counted as an error.  A track that
answers a 5Mb region with 600MB is one this robot cannot test, which is the same
situation the row count screen already catches before submitting; counting it
would put a failure in every weekly run and leave the summary as useless a gate
as the one that never failed.

The log is line buffered now as well.  Finding out that a run died partway
through is what this robot is for, and a block of buffered lines lost on the way
out is part of how the old failure left no trace of which track it was on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

diff --git src/inc/net.h src/inc/net.h
index 2e050b9246a..f3dac13932b 100644
--- src/inc/net.h
+++ src/inc/net.h
@@ -178,30 +178,35 @@
  * will skip any headers.   Free this with
  * lineFileClose(). */
 
 struct lineFile *netLineFileMayOpen(char *url);
 /* Same as netLineFileOpen, but warns and returns
  * null rather than aborting on problems. */
 
 struct lineFile *netLineFileSilentOpen(char *url);
 /* Open a lineFile on a URL.  Just return NULL without any user
  * visible warning message if there's a problem. */
 
 struct dyString *netSlurpFile(int sd);
 /* Slurp file into dynamic string and return.  Result will include http headers and
  * the like. */
 
+struct dyString *netSlurpFileMax(int sd, size_t maxSize);
+/* Slurp file into dynamic string and return.  If maxSize is nonzero and the data runs
+ * past it, stop reading, free what was read, and return NULL.  Zero means no limit,
+ * which is what netSlurpFile does. */
+
 struct dyString *netSlurpUrl(char *url);
 /* Go grab all of URL and return it as dynamic string.  Result will include http headers
  * and the like. This will errAbort if there's a problem. */
 
 char *netReadTextFileIfExists(char *url);
 /* Read entire URL and return it as a string.  URL should be text (embedded zeros will be
  * interpreted as end of string).  If the url doesn't exist or has other problems,
  * returns NULL. Does *not* include http headers. */
 
 struct lineFile *netHttpLineFileMayOpen(char *url, struct netParsedUrl **npu);
 /* Parse URL and open an HTTP socket for it but don't send a request yet. */
 
 void netHttpGet(struct lineFile *lf, struct netParsedUrl *npu,
 		boolean keepAlive);
 /* Send a GET request, possibly with Keep-Alive. */