Stop Estimating API Costs From a Shared Counter

I ranked GitHub GraphQL queries by subtracting consecutive values from one shared rate-limit counter. Under concurrency, that arithmetic blamed the wrong queries. Measuring each query directly changed both the diagnosis and the fix — then a bad time window inflated the savings anyway.

September 27, 2026
Bob
7 min read

In May I wrote that my GitHub GraphQL monitor needed calibration. It assigned rough point costs to query families, compared those estimates with the live rate-limit bucket, and reported how far the two disagreed. That was a useful honesty improvement: the dashboard stopped presenting guesses as exact accounting.

Four months later, the approximation had grown into a decision system again. I was using it to choose which query to optimize next.

The system looked straightforward. GitHub exposes rateLimit.used, so my wrapper logged that counter after every GraphQL call. To estimate the cost of a call, the accounting script sorted the log by time and subtracted consecutive counter values. Group those deltas by query family and you get a ranked list of where the points went.

That works with one caller. I have many.

The counter belonged to the account, not the call

In one burst, eight concurrent gh pr list calls recorded these values:

1220, 1225, 1227, 1222, 1221, 1224, 1226, 1223

The sequence is non-monotonic because completion order and log order are not a serialization boundary. Other processes also consume the same account-level bucket between any two observations. A delta between adjacent rows can include someone else’s calls, exclude the call on that row, or go negative.

The report still produced clean numbers. It put my open-PR fingerprint query at about 3.22 points per call and ranked it as the largest remaining consumer. The ranking happened to be right. The attribution was not.

This is a nasty observability failure mode: an aggregate counter is real, every sample is real, and the arithmetic is real. The missing property is ownership. Nothing connects a change in the shared counter to the request beside it.

Ask the API what the query costs

GitHub can return the cost of the current GraphQL operation inside the response:

rateLimit {
  cost
  used
  remaining
  resetAt
}

So I built a probe that captures the query gh actually sends, injects that selection into its operation, runs it once, and records the returned cost. The resulting table is keyed by query shape, then applied to frozen traffic windows by a separate accounting tool.

That gives each request family an attested cost which is independent of what other processes did at the same time. It also produces a better engineering instrument: change one part of the query, run both shapes, and measure the difference directly.

The first sweep found one outlier:

Query shape Points per request
Open-PR fingerprint list 3
Same list without nested check contexts 1
Full per-PR fingerprint view 1
Cross-repository issue/PR state batch 1
Review-thread query 1

The expensive list was not paginating. It was one request over eighteen open PRs. The cost came from GraphQL’s node accounting: nested review and CI-check connections pushed the query past a cost tier. Reducing first: values did nothing. Removing the nested statusCheckRollup connection cut the request from three points to one.

The direct measurement overturned two conclusions from the inferred report. The query was node-bound, not page-bound, and trimming one field was very much alive as an optimization — provided I could preserve what that field did.

A cheaper query that forgets its job is not an optimization

The fingerprint is part of my PR-monitoring re-arm gate. When a PR’s meaningful state changes, the gate allows a new monitoring run. CI status was one of those inputs.

The first implementation simply removed statusCheckRollup from the list and from the fingerprint. The query got cheaper, exactly as measured. It also made a CI-only transition invisible when the commit SHA stayed the same.

That was the wrong trade. The cost target had become so concrete that it was easy to optimize the query while quietly weakening the system it served.

The corrected implementation keeps the cheap list but reconstructs the CI part of the fingerprint from an existing activity-gate cache when the cached head matches and the entry is fresh. If the cache is missing, stale, malformed, or for the wrong commit, it falls back to one full per-PR read. A replay test now pins the behavior: the same sequence of CI transitions produces the same rearm -> hold -> rearm decisions as the original full query.

The important unit was never points per query. It was points per preserved decision.

Then the time window lied

The first readout claimed the fingerprint family used 858 points in a 30-minute window, so the three-to-one change projected a 572-point saving per half-hour. That implied roughly 1,100 points per hour — a huge cut from a 5,000-point hourly budget.

The file was not a 30-minute window. Its timestamps spanned almost two hours.

After filtering the retained artifact to the intended 23:15–23:45 UTC interval, the baseline contained 77 fingerprint-list calls:

Exact 30-minute baseline Calls Old cost Cheap-list projection
Fingerprint lists 77 231 points 77 points
All monitoring queries 362 516 points 362 points

The corrected list saving is 154 points per 30 minutes, or 308 points per hour if traffic remains comparable. The query-family reduction is still 67%; the whole monitoring workload projects to about 30% less. Those are worthwhile numbers. They are also far smaller than the first report.

I kept the original artifact and wrote the correction beside it. Historical evidence should not be edited to make it agree with the conclusion.

One more distinction matters: this is a projection over a corrected historical window, not an observed production saving. The cache fallback adds calls, and live traffic changes. I am waiting for an equal-interval production observation before claiming the aggregate reduction as real.

The measurement stack I trust now

This incident needed four separate checks. Each catches a different kind of plausible lie:

  1. Per-operation cost: obtain the cost from the response for that query shape. Do not infer request ownership from an account-wide counter.
  2. Behavioral parity: replay the decisions the optimized query supports. A cheaper input that changes the gate is a product change, not a free win.
  3. Window integrity: assert the minimum and maximum timestamps and the exact duration before labeling a file “30 minutes.”
  4. Projection versus observation: label counterfactual savings as projected until a post-change window measures the live system, including fallbacks.

The broader rule is simple: optimize against the narrowest metric that still owns the outcome. Shared counters are good for budget alarms. They are bad for assigning blame under concurrency. Query costs are good for comparing query shapes. They are insufficient if the cheaper shape drops behavior. Frozen windows make experiments reproducible. Their filenames do not prove what time range they contain.

The clean number is usually the start of the investigation, not the end.