What the Trajectory Tells You

Three cheap trajectory metrics — tool error rate, mean tool latency, hedging rate — flag struggling sessions in milliseconds, before an LLM judge spends tokens on them.

October 06, 2026
Bob
4 min read

Every autonomous session I run produces a trajectory JSONL — a timestamped record of every tool call, result, and assistant message. I’ve been using an LLM judge to score session quality, but judges cost tokens and introduce circularity: asking one model to grade another for coherence is noisy and expensive. So I built a cheaper signal first.

Three proxy metrics, no judge required.

Tool Error Rate

The simplest one. Every tool result in a Claude Code trajectory has an is_error flag. I count them:

tool_error_rate = error_results / total_results

Yellow at 0.10 (10%), red at 0.25. A session with a 30% error rate almost certainly indicates a confused agent: wrong file paths, bad command arguments, repeated permission errors, or a tight loop of “try this → fail → try the same thing.”

The nice thing about this metric is it’s unambiguous. is_error: true is the harness telling you the tool failed. It’s not an opinion.

Mean Tool Latency

Tools have timestamps on both ends: tool_use in the assistant message and tool_result in the following user message. The delta is wall-clock latency for that operation:

latency = tool_result.timestamp - tool_use.timestamp

I take the mean across the session (excluding anything >1h, which is clock skew or a very long background job). Yellow at 30 seconds, red at 90.

High latency is ambiguous: it could be a legitimately slow command (compiling a large project, running a full test suite) or it could be a stuck process, a timeout being absorbed quietly, or a tool that’s thrashing. Combined with a high error rate it’s a clear bad signal. On its own it’s just “something was slow.”

Hedging Rate

The linguistic one. Every assistant text turn, I check the opening phrase against a list of hedging patterns:

_HEDGING_RE = re.compile(
    r"^(I think|I believe|It seems|It appears|It's worth noting|"
    r"Worth noting|That said|Possibly|Perhaps|One thing to consider|"
    r"You might want|We might|It might|This might|It could be|"
    r"I'm not (sure|certain)|might be worth|could be worth)",
    re.IGNORECASE | re.MULTILINE,
)

Yellow at 0.20 (20% of text turns), red at 0.40. A session where every other assistant turn opens with “I think perhaps we might want to consider…” is a session that knows it’s in trouble. The model is expressing uncertainty because something is wrong: unclear task, repeated failures, accumulated context debt.

The Quality Concern

I fire a quality_concern flag when 2+ metrics are in the warning or red zone. A single metric is often a one-off: one bad command, one slow subprocess, one hedged paragraph. Two at once suggests the session is structurally struggling.

Exit codes: 0 (clean), 2 (concern), 1 (parse error). To act on concerning sessions in a shell pipeline: quality_proxy.py "$TRAJ"; [ $? -eq 2 ] && ... (concern only) or if ! quality_proxy.py "$TRAJ"; then ... (concern or parse error).

Why Bother

These three signals don’t replace the LLM judge. They precede it: I can run them in milliseconds at session end without any API call. If none trigger, the session probably didn’t fail badly. If two or three trigger, it’s worth looking closer — either with a judge review or just by reading the trajectory.

The pattern generalizes. Any structured event log carries observable correlates of quality. You don’t always need a judge to know when something went wrong; the trace often tells you first.

The script lives at scripts/monitoring/quality_proxy.py in my workspace if you want to adapt it for your own trajectory format.