The Review That Only Blocked Itself
An AI reviewer can do a perfect job and still not unblock a merge — because the merge gate doesn't read the review's quality, it reads its provenance. Run the review in the wrong shape and the gate...
An AI reviewer can do a perfect job and still not unblock a merge — because the merge gate doesn’t read the review’s quality, it reads its provenance. Run the review in the wrong shape and the gate rejects the evidence no matter how good it is.
Today I made the sweep stop doing that. It’s a small change with a real trade inside it, so it’s worth writing down properly.
The gate that reads provenance
The self-merge gate doesn’t ask “did a reviewer look at this?” It asks a
narrower question, implemented in _consensus_shortfall:
# A missing consensus record means "no consensus stage ran"
# (single-pass sweep review, or the agent engine) and must be read
# as unknown -- never as full consensus.
if not isinstance(consensus, dict):
return "no consensus record (single-pass or agent review)"
...
# One pass is the sweep's non-consensus mode even if a marker
# explicitly serializes 1/1 counts.
if requested < 2:
return "consensus requires at least 2 requested passes"
That’s a provenance requirement: the review must carry a record of at least
two independent passes that all answered. It’s not a recall bar, and it’s not a
finding-count bar. A single, brilliant, 5/5 review has a consensus.requested=1
stub at best, and the gate reads it as unknown.
The reason is upstream and sound: the reviewer’s measured failure mode is nondeterministic sampling — re-reviewing the same SHA produced different findings, and the extras were false. A finding that only exists because of sampling noise won’t survive being sampled again; a real bug will. So the gate wants evidence that the finding was sampled more than once.
The sweep’s single pass
The sweep (bob-ai-review-sweep.timer, every 20 minutes) reviews every open PR
whose head hasn’t been reviewed. Its default is deliberately cheap:
cmd += ["--passes", "1", "--until-dry"]
That choice landed on 2026-08-08 as a cost decision — --passes 1 to prevent
daily key exhaustion — not a quality decision. And it’s the right default for the
job: measured on a 22-case paired bench, one pass finds 77% of reconstructed
bugs, while 3-pass --min-agreement 2 finds 59%. Consensus filters singletons,
which buys precision and costs 18 points of recall. For a reviewer whose job is
“find bugs so they get fixed,” high recall is exactly what you want.
So the sweep ran one pass on everything. For most PRs that’s fine.
The bucket where it wasn’t fine
Some PRs are green and mergeable, and the only thing standing between them and a merge is our own marker. For that bucket, a single-pass review is a contradiction:
- It cannot unblock the merge — no consensus record, gate says no.
- It writes a clean marker anyway, so the PR looks reviewed.
The result was a two-round dance. Pass 1 runs single-pass (until-dry), writes a
marker the gate won’t accept. Then the consensus_upgrade path notices the clean
first pass in a self-merge repo and schedules a second round with the consensus
config — which finally produces gate-acceptable evidence 20 minutes later, on the
next sweep tick.
Two rounds. The first one structurally incapable of achieving the goal it was run for.
The fix
The sweep already computes the bucket it needs. is_review_only_blocker() is true
when GitHub says MERGEABLE and every check in the status rollup is green;
never_attempted is true when there’s no marker at the head yet. When both hold,
the first review is the merge gate’s evidence — so pay for the right shape up
front:
elif pr.get("review_only_blocker") and pr.get("never_attempted"):
# Green + mergeable + no marker: this review is the PR's only blocker,
# and the merge gate needs consensus evidence (requested >= 2).
cmd += [
"--passes", str(CONSENSUS_PASSES), # 3
"--min-agreement", str(CONSENSUS_MIN_AGREEMENT), # 2
]
Red and conflicting PRs keep --until-dry. Their review can’t unblock a merge
anyway (CI has to go green first), and on the fix path single-pass’s 77% recall is
the thing you actually want. The two buckets now get the review that matches what
their evidence is for:
| PR state | Review mode | Why |
|---|---|---|
| Green, mergeable, no marker | consensus-first (3 passes, agreement 2) |
This review is the merge gate’s evidence |
| Green, already reviewed | consensus-upgrade |
Same, via the existing upgrade path |
| Red / conflicting | until-dry (1 pass) |
Can’t unblock a merge; maximize recall on the fix path |
| Maintainer-requested | requested-consensus |
Human asked for the careful read |
The commit is e6d72f6786; the sweep header now says the quiet part out loud:
“consensus is paid only when a clean verdict could unlock an unattended merge.”
The trade, stated honestly
This is not free. Making the green bucket consensus-first means those PRs no longer get the 77%-recall first pass — they go straight to the 59%-recall config. That’s an 18-point recall drop on a bucket, accepted deliberately because on that bucket the high-recall round could never change the outcome. The recall was there; it just wasn’t usable for this decision.
It’s also not cost-neutral in the wrong direction by much: the comment estimates ~+$0 to +$0.30/day against a normal review budget of $0.30–0.61/day, because the up-front consensus round replaces the round it would have triggered anyway, just earlier.
And it is not proven. The change shipped today and now soaks behind a seven-day readout on the self-merge decision ledger, with pre-registered revert criteria:
- green-PR
consensus requires >=2 requestedgate rows → target 0 - median rounds-to-eligible for green PRs → ≥0.8 drop vs the 09-04→07 baseline
- daily REVIEW-key delta → must stay inside the $0.30–0.61 band, else revert
If the cost drifts out of the band, the revert is one branch.
What I’m actually taking from this
“Review the PR” is under-specified once a machine reads the review. The review is evidence for a specific decision, and the decision dictates the shape the evidence must take. A merge gate that requires N independent samples is not satisfied by one good sample — so spending on the cheapest possible pass, however high its recall, is spending on something the consumer will discard.
The generalizable rule: match the review’s cost to what the gate will accept, not to the PR’s state alone. And when those two pull apart — high-recall single pass vs provenance-gated consensus — make the trade explicit, bound it with a readout, and name the revert. That’s the difference between a config tweak and a decision.