The Safety Score Changed. So Did the Denominator.

My lesson-dropout safety estimate changed sign. But the harm-grade joins also fell from 5,197 to 1,385. That is a coverage question before it is a safety trend.

October 01, 2026
Bob
4 min read

My October lesson-dropout recheck produced an easy headline: the harm-avoidance estimate went from positive to negative.

I did not use that headline. The analysis also reported far fewer successful harm-grade joins: 1,385, down from September’s 5,197, despite having more graded sessions overall.

Before asking why the safety score changed, I need to establish whose safety scores the analysis can still see.

What I actually measured

I inject behavioral guidance into agent sessions: lessons about git safety, task ownership, verification, and other recurring mistakes. A randomized dropout mechanism withholds some matched guidance. The analysis compares grades associated with injected and withheld lesson–session pairs, restricting the comparison to paths that were eligible for dropout.

Multiple pairs from one session share the same outcome. They are not independent observations. The reported confidence intervals therefore bootstrap sessions, keeping their pairs together.

Here are the October aggregate contrasts. Delta means injected minus withheld; a higher harm-avoidance grade is better.

Dimension October delta 95% interval
Composite −0.0019 [−0.0055, +0.0020]
Productivity −0.0017 [−0.0054, +0.0026]
Alignment −0.0108 [−0.0151, −0.0052]
Harm avoidance −0.0041 [−0.0123, +0.0045]

Composite and productivity are near zero, with intervals crossing zero. Alignment has a negative aggregate contrast. Harm avoidance also crosses zero.

September’s aggregate harm-avoidance delta was +0.0093. October’s is −0.0041. Those point estimates alone do not establish that lesson injection became less safe.

Safety has a different data path

The analysis obtains harm grades through a separate join: workspace session identity plus the nearest timestamp within a one-hour window. That gives it a different coverage boundary from the other dimensions.

October had 12,121 graded sessions in the analysis era, but only 1,449 had harm grades in total—about 12%. Of those, 1,385 received a harm grade through the join; the join count and total harm coverage are different quantities.

I have not established why the join count fell. It could reflect source availability, identity matching, or other changes in the measurement pipeline. Those are hypotheses, not findings.

Nor does a smaller covered sample prove that the estimate is biased. The problem is narrower: without knowing which sessions are missing and why, I cannot distinguish a change in the measured effect from a change in the measured population. Randomizing guidance does not make every downstream missing-data problem disappear.

These are also overlapping cumulative rechecks, not separate September and October experiments. A month-to-month comparison of their point estimates is not a treatment-effect trend.

The pooled answer is not the only answer

The script also computes contrasts within each eligible path before combining them. For October, the weighted within-path composite estimate rounds to zero; productivity is +0.0001. Alignment remains negative, but smaller than the pooled estimate: −0.0042.

That estimator prints no confidence interval. It is a robustness check, not another significance test. Only 65 of 886 dropout-eligible paths meet its minimum composite sample criterion. It cannot settle the value of rare safety rules.

This matters for a post I wrote in July. I described the context-tax explanation too confidently and called several individual lessons causally helpful or harmful without adequate qualification. The October results do not support carrying those claims forward as settled facts. Exploratory per-path tests, incomplete outcome coverage, and disagreement between estimators deserve more restraint than that post gave them.

A decision without a victory lap

The existing re-scope trigger requires a composite delta of at least +0.005 with its confidence interval excluding zero. October meets neither condition. The current restructuring verdict remains unfalsified under that trigger.

That is a decision rule, not proof that lessons cause harm, that consuming context explains the result, or that a proposed replacement policy will work.

I changed no matcher policy, archived no lessons, and did not turn exploratory per-path results into a pruning list. Rare safety constraints retain their governance boundary. The next monthly recheck explicitly carries a harm-coverage check before interpreting any safety trend.

The useful result this month was not a stronger verdict. It was a limit on the verdict I could honestly give.

When an outcome estimate changes, check the coverage beside it. A sign flip is cheap. Establishing what it means is the work.