A one-line routing bug made private-repo reviews 14× slower

Private-repo reviews took 284 seconds because no-train routing sorted by price. Keeping the pinned provider order dropped p50 to 20 seconds without weakening the data-policy split.

October 06, 2026
Bob
5 min read

For six weeks every code review I ran against my own private repo took 280–340 seconds. Reviews against public repos took 20–30 seconds. I treated that as private-repo overhead: secret scanning, smaller quotas, something inherent.

It was a routing accident. The data-policy split was correct. The speed penalty was not.

The split that stayed, and the sort that should not have

I route reviews through OpenRouter. Public-repo traffic can go to a pinned provider order (together, fireworks, inceptron). Private-repo diffs must not reach a provider that trains on submitted data, so those requests set data_collection: deny.

That split is deliberate. Evals and tool-format traffic are welcome in training corpora; private diffs are not.

The no_train branch used to skip the policy’s provider order and fall through to sort: price. Cheapest eligible host first, inside the deny set. In a seven-day window that sent 201 of 212 local-lane requests out unpinned:

Answering provider n elapsed p50 / p90
OpenInference 94 209 / 638 s
Sail Research 80 337 / 704 s
Together (pinned) 7 30 / 94 s

One pass on the private repo took 3.5× longer than three passes on a public one. The engine was the same. The route was not.

The comment that justified the skip was real when it was written. A policy allowlist can contain a training endpoint (deepseek did, until OpenRouter dropped it on 2026-09-10). Intersecting that allowlist with data_collection: deny can empty the provider set and 404. order is a preference, not a restriction. It cannot empty the set. The hazard was gone. The sort: price branch stayed.

What a written target caught

An internal review of the in-session reviewer on 2026-10-03 scored it 3/10. Pre-commit latency on the brain repo was 284 s. That scored 1, not 0, because 284 is technically ≤300. You only see that kind of miss when the number is written down next to a target.

The same review found the rest of the delivery path:

  • 23% of wrapper sessions never collected a verdict. The reviewer ran in the background, the session timed out polling, and the result evaporated.
  • 474 of 641 ledger rows were poll rows. Each collect call appended a row instead of waiting.
  • There was no health check reading those ledgers.

Latency was the loudest symptom. Delivery was also broken.

Three fixes, one afternoon

Routing. When no_train is on and a policy order exists, keep order, set allow_fallbacks: true, and drop sort: price. Deny still filters training hosts. Price sorting remains only when there is no order to prefer. One live smoke-check after the change routed the first brain-repo request to Together in 27.8 s.

Collect. --wait blocks until the unit finishes and writes a single terminal ledger row. The still-running collect path prints a status line instead of appending another row.

Health. self-review.py --check in_session_review reads the wrapper ledger and warns on never-collected share, p50, deadline timeouts, and orphan artifacts. The 30% never-collected rate had been running for weeks without a page.

Three days later

Re-score on 2026-10-06, 163 wrapper sessions and 251 local-lane requests after the landing commits:

Criterion Before After
Verdict reached the author 70% 80.4%
Wrapper p50 (brain repo) 284 s 20.2 s
Hard-deadline share 4.2% 0.8%
Poll rows / wrapper orphans 474 / 15 0 / 0
Health observable none check exists, ran OK
Wrapper ÷ direct-engine 0.34 0.66
Total 3/10 10/10

The 14× drop, 284 s to 20.2 s, is the visible one. Private-repo reviews now feel like public-repo reviews. p90 is still 206 s; the typical case is fast, the tail is not gone.

Verdict-reached only moved to 80%, short of 85%. 32 sessions still exited background_no_artifact: the unit launched and died without writing a finding (model error or OOM). That is reviewer reliability, not routing.

Sessions still call the raw engine directly at nearly 2× the wrapper rate. The wrapper is now fast enough that more sessions choose it. The two call paths are still not one path.

The comment was true. Then it wasn’t.

The 284 s number sat in the ledger the whole time. Nothing compared it to a target, so it felt like the cost of reviewing a private repo.

The no_train comment was accurate on the day it was written. The dependency it named was removed three weeks before the fix. Nobody updated the comment. Nobody asked whether the skip was still load-bearing.

A comment that says “this must be X because of Y” should carry Y as a machine-checkable probe. Prose rots. The policy order, when present, is now the route. Price sorting is the fallback for models that have no order, which is the case the probe still covers.

The remaining stall question is not first-byte. Later measurement showed timed-out requests already had a header; a header deadline would have killed none of them. The open gate is idle time after the first byte, and it is still a measurement task, not a new deadline.