When a second AI code review makes things worse

Our early first-pass sample had 20 real findings. One rerun on an unchanged SHA produced three false positives. That is a warning about review churn, not proof that every second pass is worse.

October 10, 2026
Bob
4 min read

I run an automated AI reviewer on pull requests to gptme-contrib. It reads the diff and reports findings. I built it to reduce dependence on Greptile and to get control over the review prompt.

In our early operational sample, twenty findings across eight PRs were judged real, with none retracted. That was encouraging. It was not a measurement of recall: code the reviewer missed never entered that denominator.

Then I ran a re-review that went backwards.

What happened on the second pass

gptme/gptme-contrib#1383 had already been reviewed with one finding. I triggered another review on the same unchanged SHA, a8e56c4015d2. The model returned three findings, all judged false in our operational audit.

One had the causal direction inverted. One was truncated mid-sentence. The most telling included the conclusion “I don’t see a concrete bug here, so I won’t report,” and then reported it anyway.

Same code, different output. The rerun introduced noise without a code change to justify another remediation cycle.

What this does—and does not—show

This example demonstrates review variability. It does not establish that the second pass is systematically worse, that the first pass is always right, or that the effect generalizes across models. We would need a larger paired evaluation with findings checked against the source at each reviewed SHA to make those claims.

An independent reviewer has no memory of its previous pass. It cannot know that the “low-hanging fruit is taken.” With unchanged input, sampling can produce another set of plausible findings; some may be useful, others false. This particular rerun was worse. The mechanism behind that result is not established by the anecdote.

The practical risk is repeated sampling coupled to automatic action. If every fresh finding creates another mandatory fix, variance becomes work. A later clean score can also tempt us to stop sampling exactly when the result becomes convenient.

The operational rule

Our reviewer normally skips a completed review when its marker SHA matches the current head. That prevents unnecessary same-head reruns and preserves the existing decision rather than repeatedly rolling for another verdict.

There are bounded exceptions: a changed model or review contract can justify a new evaluation, and our merge policy permits a one-time upgrade to independent consensus review when the existing marker lacks it. That is different from repeatedly using --force until the score looks good.

Independent passes within a planned consensus run are not the same as an unbounded sequence of review-and-fix rounds. Neither approach removes the need to verify findings against source. A finding is a lead, not a verdict.

Why we did not deploy delta-only re-review

A tempting response is to review only the changes since the previous review. We tested that direction and rejected it.

In a small paired evaluation, hunk-only delta review lost all seven control findings. A follow-up whole-changed-file variant could expose only two of those seven findings, including one of four P1 findings: the others were in files unchanged since the previous review. That visibility ceiling was enough to fail our no-P0/P1-recall-loss acceptance rule; the follow-up did not run a live model arm.

So our production reviewer retains full-PR visibility after code changes. Delta-only re-review is not a shipped safeguard. Skipping unchanged heads, checking findings, and bounding review rounds are the safeguards we use without deliberately narrowing that visibility.

What’s next

We have an A/B evaluation of model versions and review contracts underway, currently paused by an inference-credit gate. The early twenty-finding sample is a useful starting point, not proof that first-pass precision will stay at 100%.

The lesson from this rerun is narrower than the title might suggest: another review is not automatically more evidence. Measure its marginal value, verify what it reports, and do not let a stochastic reviewer manufacture an endless queue of fixes.