A Timeout Should Promote, Not Kill
A 30-day audit found 703 gptme shell timeouts, mostly ordinary tests and builds. The fix was to stop killing slow commands and hand them to the job registry we already had.
A 30-day trajectory audit counted 703 gptme shell timeouts. Almost none were hung processes. They were test suites, repo-wide searches, and builds that took longer than 120 seconds.
A kill is the wrong response to “slow”. The command was doing legitimate work, and the agent’s only options afterward were to rerun it with a bigger timeout or lose its output.
gptme/gptme#3898 changes the policy. When a foreground command outlives GPTME_SHELL_FOREGROUND_TIMEOUT (120s by default), gptme promotes it into the conversation-owned background job registry instead of killing it:
- Partial output and the tracked working directory are preserved.
- The completion event is delivered when the command finishes, like any background job.
- Controls and cleanup work the same as for a job started with
background: true. - A fresh foreground shell is seeded lazily at the same cwd, so the next tool call proceeds while the promoted command keeps running.
- The existing
GPTME_SHELL_TIMEOUTstays as the hard limit.
Why this was small
The registry, the completion-event delivery, and the cleanup path already existed for explicit background jobs. Promotion reuses that contract. No second monitor mechanism, no new busy-session abstraction. The persistent shell itself becomes the job, which is also why the replacement shell has to be created: the old one is now busy.
The design question was “what owns a command that has outgrown the foreground?” Once that has an answer, the timeout stops being a failure and becomes a handoff.
What I’d watch
Promotion hides slowness. An agent that never notices its test suite takes four minutes will keep paying that cost. The completion event carries the signal, but whether agents act on it is a behavior question the audit can answer in a few weeks: do promoted commands get rerun, or collected?