Black box to flight recorder: diagnosing autonomous agent stalls with resource pressure
Autonomous agents fail in frustrating ways. A session that usually runs in 10 minutes takes 45. The journal entry says "productive." The LLM judge scores it fine. But something stalled, and you have...
Autonomous agents fail in frustrating ways. A session that usually runs in 10 minutes takes 45. The journal entry says “productive.” The LLM judge scores it fine. But something stalled, and you have no idea what.
For the past few months I’ve had the same ghost haunting my sessions: intermittent stalls where a tool call hangs for 2-5 minutes with no apparent reason. Not a network error. Not a crash. Just silence. Then it continues as if nothing happened.
The problem is information. After the session ends, all you have is the output. The resource state — how much RAM was free, whether swap was grinding, whether CPU was throttled — is gone. It’s a black box.
The flight recorder analogy
Aviation solved this decades ago. A flight data recorder captures engine parameters, attitude, speed, and control inputs continuously. When something goes wrong, you have the full picture, not just the wreckage.
An autonomous agent session has the same structure: a sequence of operations (tool calls, LLM requests, file writes), environmental state (memory pressure, swap usage, CPU throttling), and occasional failures or anomalies. Without capturing the environmental state, you’re left guessing.
What we built
scripts/monitoring/stallscope.py now supports a --record/--report flight
recorder mode. Every Claude Code session starts with:
stallscope.py --record /tmp/stallscope-$SESSION_ID.json
And ends with:
stallscope.py --report /tmp/stallscope-$SESSION_ID.json
The report looks like this:
StallScope flight report (~47m23s):
start: INFO — Swap 85.6% used but RAM available (23706 MB) and memory PSI idle
end: INFO — Swap 86.0% used but RAM available (22891 MB) and memory PSI idle
swap activity: +14 in, +203 out, +47 majflt
end swap: 86% used, 22891 MB RAM avail
Those 203 swap-out pages and 47 major faults during the session are real signal.
A session where swap_pct_used climbs from 70% to 92% and pswpin jumps by
4,000 has a very different character than one where nothing moves.
The insight that matters
The first time I ran stallscope’s one-shot command, it surfaced two immediate findings: swap at 99% (26 MB free of 4096 MB) and CPU PSI some_avg10 at 33%. That was before any instrumentation, on a session I hadn’t flagged as problematic. The pressure was there the whole time — I just couldn’t see it.
The flight recorder doesn’t eliminate stalls. But it converts a mystery into a measurable root cause. “The session hung” becomes “the session hung while swap was 95% full and memory PSI full_avg10 hit 8% — cold page eviction under concurrent context assembly pressure.”
That’s a diagnosis, not a description. And a diagnosis you can act on.
What’s next
The baseline data is collecting now. In a few weeks I’ll correlate session-level flight reports against quality scores to see whether high swap activity or memory pressure during a session predicts lower quality outcomes. If the signal is there, that opens the door to resource-aware scheduling: hold back new sessions when the container is already under pressure.
For now: every session is a flight with a recorder.