Repository navigation
Tags: redis/redis
Tags
Explain a test [TIMEOUT] instead of killing it silently (#15879) A hung test run tells us almost nothing today. When no client has made progress for --timeout seconds, test_server_cron prints the clients' last reported state, SIGKILLs every server via force_kill_all_servers and exits; --dump-logs only fires for a failed or excepted test, never for a timeout. So a 20-minute hang costs a whole CI run and produces a few lines. Collect the evidence before tearing the run down: Crash-report the surviving servers. SIGSEGV makes redis log a stack trace of every one of its threads plus INFO, the client list and the config (printCrashReport) and then die. The handler runs on whichever thread takes the signal, so it works on a server whose event loop is wedged -- exactly the case we cannot diagnose from outside. kill_server already resorts to SIGSEGV for the same reason when a server won't exit, but the timeout path never reaches it. Then print each server's crash report, starting 10 lines above "REDIS BUG REPORT START" for context (or the log's tail if there is no report, since that is then the only evidence). Servers are children of the stuck client, which isn't reaping them, so a dead one is a zombie that kill -0 still reports alive; wait on is_running (via ps) instead. Report where in the test each client stopped, not just its last state. Crash-reporting the servers usually unblocks a client by itself -- its connection dies, the error unwinds, and the client's existing top-level handler reports $::errorInfo, a Tcl stack trace naming the exact line -- so collect that first. A client stuck on something else is poked with SIGUSR1, which it turns into a Tcl error with Tclx's "signal error": that interrupts a blocking read, a long "after" and a polling loop alike. Tclx is optional; without it a timeout simply reports no client stack trace. Order matters here: the servers must be collected first, because unblocking a client makes start_server kill the very servers we wanted a report from. Also: read_from_test_client threw "expected non-negative integer" once a reporting client exited, because we now pump the event loop while it does. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Low Risk** > Changes are limited to the Tcl test harness timeout path; no production server or runtime behavior is affected. > > **Overview** > When the suite hits **`--timeout`** (no client progress), it no longer tears down immediately after printing each client’s last task. The test server **collects diagnostics first**, then kills clients and servers as before. > > **Server evidence:** Still-running Redis instances from `::active_servers` get **SIGCONT** (if stopped) and **SIGSEGV** so they write a full crash report even when the event loop is wedged. Logs are located under `tests/tmp` by pid, then **`dump_crash_report`** prints the tail around `REDIS BUG REPORT START` (or the last 256KB if there is no report). **`is_running`** treats zombies as dead so waits don’t hang on unreaped children. > > **Client evidence:** Clients send their OS pid on `ready` and optionally advertise **`sigusr1-trace`** when Tclx is available. After server dumps, the server waits briefly for natural **`exception`**/`err` unwinds, then sends **SIGUSR1** to remaining clients so Tclx turns it into a stack trace. **`::in_timeout_report`** suppresses re-entrant timeout cron, avoids fatal handling of those packets, and fixes **`read_from_test_client`** when a client disconnects mid-report (invalid length no longer spins or crashes the handler). > > **Order:** Server crash collection runs **before** client stack traces so unblocking a client doesn’t tear down servers before their reports are captured. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 5927e7e. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY --> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Add aof_cmd_duration estimate for AOF reload RTO visibility Expose a best-effort AOF replay-time estimate in INFO persistence so operators can gauge reload RTO. Count time only for writes that enter the AOF, credit leftover call() time to synthetic rewrites, and skip the bookkeeping when AOF is off.
PreviousNext