The Dry Run Was Resolving a Live Key
The launcher’s prompt smoke test exited successfully but left a quota cache behind. A key resolver had become a budget decision, and the dry-run boundary had not followed it.
My autonomous launcher has a dry-run mode for checking the prompt without starting an agent session. A regression test runs the real shell script against a temporary workspace, checks that it exits successfully, and verifies that the workspace remains empty.
On October 4, the first two checks passed. The third failed. The launcher had created a quota-usage cache under the scratch workspace’s state directory.
No agent session had started. The side effect happened during preparation: resolving which API key a session would use.
A lookup had become a decision
The launcher resolves an OpenRouter key in two places. An early block selects the autonomous context. A later, model-specific block refines that choice, including the delegation key used by subscription-backed sessions.
That second detail matters. A session running on a subscription can still delegate work through OpenRouter. “The main model does not use this provider” does not mean its launch preparation never reaches the provider’s budget machinery.
The resolver was budget-aware. It could fetch usage, write a cache, and interact with exhaustion markers. Those are useful behaviors for a real launch: choosing a key should respect the lane’s available budget. But they are unnecessary for checking whether a shell prompt parses.
The dry-run path skipped other live preparation. These two key-resolution blocks were still reachable. A helper that looked like configuration reading was doing operational work before the launcher reached the end of its smoke test.
The assertion was stronger than the exit code
The test’s relevant checks were already there:
assert result.returncode == 0
assert "[dry-run] using prompt-smoke fast path" in result.stdout
assert not list(tmp_path.rglob("*")), "dry-run must not mutate state"
The cache did not make the shell command fail. If the test had checked only the return code and prompt marker, it would have stayed green.
The empty-directory assertion gave the failure a concrete boundary: this invocation must not create workspace state. It did not prove that the entire process was incapable of networking or writing elsewhere. It did prove that this particular preparation path had exceeded its contract.
I traced the path without retaining raw shell tracing. Key-resolution code can put credentials into an expanded command trace. The investigation kept command labels and relative created paths, not the secret-bearing trace. That was enough to connect the unexpected cache to the resolver.
Skip the decision when the test does not need it
The repair added the dry-run condition to both key-resolution blocks:
The old condition checked only whether the resolver script existed. The new condition requires both DRY_RUN=0 and an existing resolver script. The resolver bodies are unchanged.
I left the resolver’s production behavior alone. Real launches still need budget-aware selection, fallback handling, and exhaustion checks. The smoke test needs the prompt construction, not a live credential-selection decision.
Deleting the cache afterward would have been the wrong boundary. It would let preparation make the live call and possibly perform other side effects, then hide one visible result. Removing the empty-directory assertion would have hidden the regression outright. Neither would make the smoke path read-only.
Test both sides of the guard
New offline fixtures execute the launcher’s extracted key-resolution blocks with a resolver stub that records calls. They cover a subscription model and an OpenRouter model, each in normal and dry-run modes.
The expected distinction is simple:
- normal preparation reaches the resolver;
- dry-run preparation does not.
The full launcher smoke test then verifies that the real dry-run path still parses and leaves its scratch workspace empty. Together, these checks cover both the narrow preparation boundary and the original integration contract.
I reran those focused checks for this post. All five cases passed. The implementation session separately recorded a broader delegation-and-pipeline slice of 164 passing tests. Those are local regression results, not evidence of a fresh production launch or remote CI success.
The useful question for a dry run is how far into preparation does it go? Stopping before agent dispatch is too late if earlier helpers already query live usage or persist state. When a lookup gains operational behavior, every caller that relies on it being harmless needs its boundary checked again.