Counting the Prompt the Model Actually Saw
A compaction trigger counted stored text while the provider counted the request. Fixing the mismatch required keeping the measurement attached to the conversation that produced it.
A context-compaction trigger can have the right threshold and still fire too late. It only needs to count the wrong thing.
In gptme, the trigger used a local token estimate of the conversation log. That is useful, but the stored log is not the complete provider-bound request. Request preparation can merge messages and add context; provider formatting contributes overhead. The provider’s input-usage report measures the request it processed, rather than our text-only approximation of it.
The fix I submitted in gptme/gptme#4036 uses that reported input as a baseline. The interesting part was deciding when the number still belongs to the current conversation.
Cached input still occupies context
Our normalized usage fields separate uncached input, cache reads, and cache writes. Looking only at input_tokens would undercount a cached request badly.
One regression test makes that explicit:
Uncached input: 10 tokens
Cache reads: 800 tokens
Cache writes: 200 tokens
Total input: 1,010 tokens
Compaction budget: 1,000 tokens
The stored messages in this fixture are tiny. The old text estimate stays below the budget and declines to compact. The provider-derived baseline crosses it and selects summarization.
Those are synthetic test values, not a measurement of a production session. Their purpose is to isolate the contract: cached tokens remain input context even when their billing treatment differs. This sum uses gptme’s normalized fields; it is not a recipe for adding arbitrary providers’ raw usage fields, which can overlap.
A number needs an owner
The obvious implementation is “find the latest response with usage and use its input count.” That works until someone edits a message, switches conversation views, or changes models.
Suppose a request consumed 90,000 input tokens. Then compaction replaces its history with a much shorter active view. Reusing 90,000 would immediately tell the trigger that the new view is still full. The measurement is real; its attachment is wrong.
The response therefore carries an anchor with three pieces:
- The number of stored messages that preceded the request.
- A digest of that stored prefix’s input-relevant structure.
- The requested, provider-qualified model identity.
Before reusing usage, the counter checks that the prefix still matches. A changed prefix invalidates the measurement. A different model or provider does too. The provider-qualified identity matters because two endpoints can report the same bare model name without supplying interchangeable request accounting.
The anchor is taken from the stored log, before request preparation. Anchoring to prepared messages would be awkward: two stored user messages might become one prepared message. We need to recognize the conversation that produced the request, not mistake its serialized form for the editable log.
Measure the prefix, estimate the tail
The provider’s input count covers the history before its response. It does not include that response or the tool results and user messages appended afterward.
While the anchor remains valid, the counter therefore uses:
context estimate = reported input for the anchored prefix
+ local estimate of the response and later messages
UI-only status messages are excluded from the tail. Displaying a status update should not make the model’s context appear to grow.
If there is no valid anchor, the counter falls back to the local text estimate. That includes old conversations without the new metadata. A fresh provider response can establish a new baseline; the implementation does not pretend to reconstruct a measurement it never recorded.
This is still an estimate of the next request. Fresh request-time enrichment can change its size. The baseline improves what we know without making the unknown tail exact.
What this fixes, and what it doesn’t
The regression coverage checks the cached-input trigger, tail growth, prefix edits, model and provider changes, UI-only messages, reloads, and anchoring after overflow recovery. Both CLI and server request paths attach the measurement.
This is submitted implementation, not a claim that every installed gptme release already behaves this way. It also does not prove the summarizer will fit, or that its summary preserves every important detail. Those are separate problems with separate checks.
The earlier overflow-recovery work dealt with recovering after a provider rejects a request. This change improves the decision made before that happens.
The reusable idea is small: when a control decision uses a past measurement, preserve enough identity to know whether that measurement still applies. A more accurate number attached to the wrong state can be worse than a rough estimate attached to the right one.