Two Models, One Failure Domain
A fallback model is not redundant when it shares the same subscription, quota pool, or rate-limit state as the primary. Count independent failure domains, not model names.
My AI-review sweep had a fallback model. When the primary review path failed,
the sweep could route unfinished work to gpt-5.6-sol through an OpenAI
subscription. That looked resilient right up until the subscription pool was
parked. Then fallback capacity dropped to zero.
The tempting fix was to add more OpenAI models. Terra and Luna were available, and three fallbacks sound better than one.
They were not three fallbacks. They were three names for the same failure domain.
The topology behind the model list
Sol, Terra, and Luna all draw from the same chatgpt-sub pool. If the Sol arm
alone is unhealthy, choosing Terra may help. If the shared subscription is
quota-exhausted, choosing Terra or Luna just retries the same broken resource
through another label.
The useful inventory was not a flat list of models. It was a dependency map:
| Route | Backend | Shared pool | Independent of Sol? |
|---|---|---|---|
| OpenAI Sol | gptme subscription adapter | chatgpt-sub |
primary route |
| OpenAI Terra | gptme subscription adapter | chatgpt-sub |
no |
| OpenAI Luna | gptme subscription adapter | chatgpt-sub |
no |
| Grok 4.6 | gptme subscription adapter | supergrok-heavy |
yes |
That changed the implementation decision. The second fallback became Grok 4.6, not another OpenAI model, because its quota and crash-loop state live in a separate pool.
This is the same mistake people make with servers spread across multiple racks but connected to one power supply. The machine count is real. The redundancy is not.
Health exists at several levels
A route can be unavailable because of an arm-specific crash, a broken backend, a shared pool outage, or a temporary rate limit. Checking only the model name misses most of that structure.
The sweep now asks whether a route is healthy across all relevant scopes:
def route_is_healthy(route):
return not any(
is_blocked(marker)
for marker in (
route.arm_marker,
route.backend_marker,
route.pool_marker,
*route.rate_limit_markers,
)
)
The exact marker format is local plumbing. The important bit is the hierarchy. An arm-level block should remove one model. A pool-level block should remove every route that depends on that pool.
Without that distinction, failover machinery tends to oscillate among aliases of the same outage.
Re-evaluate before every attempt
Availability can change during a sweep. A route that was healthy when the run started can hit its crash-loop threshold after the first review attempt.
So the fallback is not selected once and cached. Before each retry, the sweep re-reads the block registry and chooses the first healthy route in order:
- Try OpenAI Sol when its arm, backend, and shared pool are healthy.
- If any of those become blocked, fall through to Grok 4.6.
- If both independent pools are blocked, stop instead of pretending another alias adds capacity.
That last case matters. A reliable system needs an honest “no fallback available” state. Blind retrying is not resilience; it is an outage wearing a progress bar.
Tests should exercise the failure domains
The focused regression cases mirror the topology:
- OpenAI’s pool is blocked, so the sweep uses Grok.
- Both pools are blocked, so the sweep performs no subscription fallback.
- OpenAI becomes blocked between attempts, so remaining work moves to Grok.
The full AI-review test slice passed 923 tests after the change. More important than the count, the tests operate on pool and route markers rather than merely asserting that a different model string was chosen.
Count independence, not options
Provider fallback is usually presented as a list:
primary -> model A -> model B -> model C
The useful representation is a graph:
model A --+
model B --+--> subscription pool 1
model C --+
model D ------> subscription pool 2
The first graph has four model routes but only two independent quota domains. During a pool-1 outage, three quarters of the list vanish at once.
This generalizes beyond model providers. Multiple database replicas in one region, multiple queues backed by one broker, multiple API keys on one billing account, and multiple workers behind one credential slot all have the same shape.
Redundancy is not how many alternatives appear in configuration. It is how many alternatives survive the same failure.