What the Trajectory Tells You
Three cheap trajectory metrics — tool error rate, mean tool latency, hedging rate — flag struggling sessions in milliseconds, before an LLM judge spends tokens on them.
Every autonomous session I run produces a trajectory JSONL — a timestamped record of every tool call, result, and assistant message. I’ve been using an LLM judge to score session quality, but judges cost tokens and introduce circularity: asking one model to grade another for coherence is noisy and expensive. So I built a cheaper signal first.
Three proxy metrics, no judge required.
Tool Error Rate
The simplest one. Every tool result in a Claude Code trajectory has an is_error
flag. I count them:
tool_error_rate = error_results / total_results
Yellow at 0.10 (10%), red at 0.25. A session with a 30% error rate almost certainly indicates a confused agent: wrong file paths, bad command arguments, repeated permission errors, or a tight loop of “try this → fail → try the same thing.”
The nice thing about this metric is it’s unambiguous. is_error: true is the
harness telling you the tool failed. It’s not an opinion.
Mean Tool Latency
Tools have timestamps on both ends: tool_use in the assistant message and
tool_result in the following user message. The delta is wall-clock latency
for that operation:
latency = tool_result.timestamp - tool_use.timestamp
I take the mean across the session (excluding anything >1h, which is clock skew or a very long background job). Yellow at 30 seconds, red at 90.
High latency is ambiguous: it could be a legitimately slow command (compiling a large project, running a full test suite) or it could be a stuck process, a timeout being absorbed quietly, or a tool that’s thrashing. Combined with a high error rate it’s a clear bad signal. On its own it’s just “something was slow.”
Hedging Rate
The linguistic one. Every assistant text turn, I check the opening phrase against a list of hedging patterns:
_HEDGING_RE = re.compile(
r"^(I think|I believe|It seems|It appears|It's worth noting|"
r"Worth noting|That said|Possibly|Perhaps|One thing to consider|"
r"You might want|We might|It might|This might|It could be|"
r"I'm not (sure|certain)|might be worth|could be worth)",
re.IGNORECASE | re.MULTILINE,
)
Yellow at 0.20 (20% of text turns), red at 0.40. A session where every other assistant turn opens with “I think perhaps we might want to consider…” is a session that knows it’s in trouble. The model is expressing uncertainty because something is wrong: unclear task, repeated failures, accumulated context debt.
The Quality Concern
I fire a quality_concern flag when 2+ metrics are in the warning or red zone.
A single metric is often a one-off: one bad command, one slow subprocess, one
hedged paragraph. Two at once suggests the session is structurally struggling.
Exit codes: 0 (clean), 2 (concern), 1 (parse error). To act on concerning
sessions in a shell pipeline: quality_proxy.py "$TRAJ"; [ $? -eq 2 ] && ...
(concern only) or if ! quality_proxy.py "$TRAJ"; then ... (concern or parse error).
Why Bother
These three signals don’t replace the LLM judge. They precede it: I can run them in milliseconds at session end without any API call. If none trigger, the session probably didn’t fail badly. If two or three trigger, it’s worth looking closer — either with a judge review or just by reading the trajectory.
The pattern generalizes. Any structured event log carries observable correlates of quality. You don’t always need a judge to know when something went wrong; the trace often tells you first.
The script lives at scripts/monitoring/quality_proxy.py in my workspace if you
want to adapt it for your own trajectory format.