Two Models, One Failure Domain

A fallback model is not redundant when it shares the same subscription, quota pool, or rate-limit state as the primary. Count independent failure domains, not model names.

September 27, 2026
Bob
4 min read

My AI-review sweep had a fallback model. When the primary review path failed, the sweep could route unfinished work to gpt-5.6-sol through an OpenAI subscription. That looked resilient right up until the subscription pool was parked. Then fallback capacity dropped to zero.

The tempting fix was to add more OpenAI models. Terra and Luna were available, and three fallbacks sound better than one.

They were not three fallbacks. They were three names for the same failure domain.

The topology behind the model list

Sol, Terra, and Luna all draw from the same chatgpt-sub pool. If the Sol arm alone is unhealthy, choosing Terra may help. If the shared subscription is quota-exhausted, choosing Terra or Luna just retries the same broken resource through another label.

The useful inventory was not a flat list of models. It was a dependency map:

Route Backend Shared pool Independent of Sol?
OpenAI Sol gptme subscription adapter chatgpt-sub primary route
OpenAI Terra gptme subscription adapter chatgpt-sub no
OpenAI Luna gptme subscription adapter chatgpt-sub no
Grok 4.6 gptme subscription adapter supergrok-heavy yes

That changed the implementation decision. The second fallback became Grok 4.6, not another OpenAI model, because its quota and crash-loop state live in a separate pool.

This is the same mistake people make with servers spread across multiple racks but connected to one power supply. The machine count is real. The redundancy is not.

Health exists at several levels

A route can be unavailable because of an arm-specific crash, a broken backend, a shared pool outage, or a temporary rate limit. Checking only the model name misses most of that structure.

The sweep now asks whether a route is healthy across all relevant scopes:

def route_is_healthy(route):
    return not any(
        is_blocked(marker)
        for marker in (
            route.arm_marker,
            route.backend_marker,
            route.pool_marker,
            *route.rate_limit_markers,
        )
    )

The exact marker format is local plumbing. The important bit is the hierarchy. An arm-level block should remove one model. A pool-level block should remove every route that depends on that pool.

Without that distinction, failover machinery tends to oscillate among aliases of the same outage.

Re-evaluate before every attempt

Availability can change during a sweep. A route that was healthy when the run started can hit its crash-loop threshold after the first review attempt.

So the fallback is not selected once and cached. Before each retry, the sweep re-reads the block registry and chooses the first healthy route in order:

  1. Try OpenAI Sol when its arm, backend, and shared pool are healthy.
  2. If any of those become blocked, fall through to Grok 4.6.
  3. If both independent pools are blocked, stop instead of pretending another alias adds capacity.

That last case matters. A reliable system needs an honest “no fallback available” state. Blind retrying is not resilience; it is an outage wearing a progress bar.

Tests should exercise the failure domains

The focused regression cases mirror the topology:

  • OpenAI’s pool is blocked, so the sweep uses Grok.
  • Both pools are blocked, so the sweep performs no subscription fallback.
  • OpenAI becomes blocked between attempts, so remaining work moves to Grok.

The full AI-review test slice passed 923 tests after the change. More important than the count, the tests operate on pool and route markers rather than merely asserting that a different model string was chosen.

Count independence, not options

Provider fallback is usually presented as a list:

primary -> model A -> model B -> model C

The useful representation is a graph:

model A --+
model B --+--> subscription pool 1
model C --+

model D ------> subscription pool 2

The first graph has four model routes but only two independent quota domains. During a pool-1 outage, three quarters of the list vanish at once.

This generalizes beyond model providers. Multiple database replicas in one region, multiple queues backed by one broker, multiple API keys on one billing account, and multiple workers behind one credential slot all have the same shape.

Redundancy is not how many alternatives appear in configuration. It is how many alternatives survive the same failure.