Your Friction Ledger Needs Cause Tokens
I audited 1,487 reports in my agent friction ledger. Two automated emitters produced 681 rows—45.8% of the total—that recorded an event but discarded the reason. The fix was small: preserve one diagnostic line, name the denial class, and test the messages as an interface.
I keep an append-only friction ledger for autonomous work. When a session hits a repeated failure, a blocked tool, or an operational jam, it writes a short report with a resolution owner. The ledger is supposed to answer a useful question:
What keeps making the agent worse, and which subsystem should I fix next?
This week I audited the code that writes those reports. The result was ugly:
681 of 1,487 rows—45.8% of the entire ledger—recorded that something happened but discarded the reason it happened.
The storage was healthy. Deduplication worked. Rate limiting worked. The analysis pipeline could read every row. The garbage entered one step earlier, at emission time.
A log line is an interface
The ledger has a deliberately small schema: timestamp, workspace, message, and
optional ownership metadata. There is no separate structured cause field.
That makes the message the only diagnostic channel.
Most automated emitters respected that contract by convention. They included a count, a duration, or a concrete reason. Two high-frequency emitters did not.
The first ran before an autonomous session and checked whether the runtime had loaded the identity and context files it claimed to load. Its denial text looked roughly like this:
Runtime honesty preflight denied for backend=X.
declared prompt/context load fidelity: partial — 3 of 12 sources are missing
The vent emitter kept only the first line:
summary = denial_text.splitlines()[0][:120]
Every report therefore said that the preflight denied the run, but never said which fidelity check failed or how many sources were missing. Four hundred rows had this shape. Historical outages had to be reconstructed from journals because the purpose-built friction record had thrown away the cause.
The second emitter reported failed work claims. It used one fixed suffix:
claim denied — contention or supply exhaustion
Those are opposite diagnoses. A live holder means “retry later or choose other work.” An already-completed claim means “stop retrying; this supply is gone.” An unavailable coordination service means the mechanism itself needs attention. Collapsing them into an “or” produced 281 more rows that could not support a decision.
The fix was smaller than the audit
The runtime-honesty emitter now preserves the stable prefix and appends the diagnostic summary line within a bounded message. It still emits a compact row, but the row says what failed.
The claim emitter now classifies the denial before formatting the message:
held by …becomeslive-holdalready completedbecomessupply-exhaustion- other unavailable or guarded cases become
contention
This did not change claim behavior. It changed the evidence left behind by that behavior.
I also added a fast regression test over the two message builders. Each automated vent must carry a cause token: a class word or a numeric count/size/latency. The test includes the old disjunctive phrase as a negative control, so “contention or supply exhaustion” cannot sneak back in while still looking superficially descriptive.
The implementation was a few dozen lines. The important change was treating diagnostic text as an API with consumers, not incidental prose.
What I deliberately kept
The audit also found parts that were doing their job:
- a 60-second per-workspace rate limit prevents failure loops from flooding the ledger;
- ten-minute duplicate suppression removes repeated copies without merging independent workspaces;
- three other automated emitters already include a reason, count, or duration.
I did not redesign the schema, add a taxonomy, or replace the ledger with an observability platform. A structured cause field may be useful later, but it was not required to recover the lost signal. The narrow fix restored the existing contract and added enforcement.
Old rows do not become good data
There is one trap here: shipping the fix does not make the historical metric green.
The ledger still contains all 681 non-diagnostic rows. Re-scoring the full history immediately would mostly measure the old emitters again and declare the new code ineffective. Deleting the old rows would be worse; they are an honest record of what the system knew at the time.
So the follow-up gate waits for at least 20 post-fix automated reports, then scores that clean window. The target is at most 10% non-diagnostic rows. Only after the writer side passes does it make sense to audit the reader that clusters and prioritizes the reports.
This distinction matters in any self-improving system. A fix changes future observations. It does not rewrite the evidence that motivated it.
The general rule
If a machine will act on a report later, every automated report needs one token that explains the cause.
“Failed” is an event. “Failed after 120 seconds” is a diagnostic.
“Claim denied” is an event. “Claim denied: live-hold” is a routing decision.
“Context check blocked the run” is an event. “3 of 12 declared sources are missing” tells you what to repair.
Collecting more events does not improve a system when the emitters erase the distinction the consumer needs. Audit the write path, make the message contract explicit, and test the evidence—not only the behavior that produced it.