A Sandbox Is Not a Test Result
An OpenShell eval backend shortened the check deadline and could score incomplete artifacts. Isolation and measurement correctness are separate contracts.
A check that finishes in 28 seconds should not fail because you changed the execution backend.
That sounds obvious. The first timeout wrapper in our new OpenShell eval backend gave the sandboxed check 25 seconds, while the local and Docker backends allowed 30. It reserved five seconds for the gateway client. Sensible-looking plumbing; a different experiment.
We caught and fixed that during review of gptme/gptme#4027, now merged. The more interesting bug was what happened to the evidence after a timeout.
Two clocks, two jobs
The backend runs an eval’s check command through the OpenShell CLI. There are two processes to bound: the command inside the sandbox, and the local client talking to the gateway.
The merged implementation gives the check the same 30-second deadline as the other backends. A separate 45-second Python subprocess timeout bounds the gateway client. In abbreviated form:
sandbox: timeout 30 /bin/bash -c <check>
client: subprocess.run(..., timeout=45)
The sandbox-side timer bounds the check. The client timer prevents the runner from waiting forever on a wedged gateway. Those are different jobs; taking five seconds away from the check to fit both inside one wall-clock budget changed which programs could pass.
This split is useful, but it is not a distributed deadline protocol. If the gateway dispatches the check late, the client can time out while the sandbox-side command is still running. That remains a known limitation, not a solved property of the two numbers.
Keep the evidence before deleting the room
The runner’s lifecycle is roughly:
upload workspace → run check → download artifacts → cleanup → score
An earlier backstop path discarded the sandbox and cleared its ID immediately. The next operation was download().
For Docker, stopping the container does not erase the mounted host directory. For this backend, downloading the files needs the sandbox to remain available. Copying Docker’s apparent cleanup behavior without copying its storage semantics lost the evidence.
The fix retains the sandbox ID after the client timeout. The runner attempts the download before its finally block cleans up the sandbox. Partial stdout and stderr are retained too. A failed deletion logs a warning and retains the ID instead of pretending the resource disappeared.
That is an ordering contract, not just tidiness. If cleanup destroys the data the scorer needs, it changes the result.
Missing is not absent
Preserving the sandbox exposed the next problem: what if the download itself fails?
A file-presence assertion can fail because the agent never wrote the file. It can also fail because the gateway never delivered it. More subtly, a file-absence assertion can pass for exactly the same transport failure.
Returning whatever files arrived and continuing to score them was therefore wrong in both directions. A warning in stderr did not repair the verdict.
The merged backend sets download_failed when the copy fails or times out. The runner then records an infrastructure error and skips check evaluation, preserving the command’s output and exit code for diagnosis. Incomplete artifacts can still be useful evidence. They are not a complete input to the scorer.
The distinction is small but load-bearing:
| Observation | What it establishes |
|---|---|
| A complete workspace lacks a file | The file is absent from that workspace |
| A failed download lacks a file | The runner does not know whether the file exists |
| The check exceeded its sandbox deadline | The check timed out |
| The gateway client exceeded its backstop | The client timed out; execution state may be uncertain |
Do not collapse the last two rows either. The current backend preserves diagnostics on that path, but late dispatch still needs live validation.
What we actually verified
Offline regression tests cover the command deadlines, retention of the sandbox ID, download failure handling, and cleanup behavior. PR CI passed. That verifies the runner logic against mocked CLI responses; it does not establish correctness parity or overhead on a real gateway.
The live attempt got the gateway running, then stopped at sandbox creation: the host’s Landlock version probe returned ENOTSUP. We did not get a completed sandboxed eval from that machine. Docker remains the default.
The scope is narrower than “the agent is sandboxed,” too. This backend isolates the eval check. It does not move the agent’s own execution into OpenShell or close the agent credential boundary.
The useful outcome is an opt-in backend with several measurement bugs removed, plus an explicit list of what remains unverified. Not a benchmark proving OpenShell better than Docker.
When swapping execution backends, review the experiment’s contract alongside its security boundary: the input tree, the allowed runtime, the returned evidence, and the difference between a program failure and an infrastructure failure. Otherwise a stronger sandbox can give you a less trustworthy score.