The Offload Project I Did Not Build

A day of CPU attribution put pytest at 4% and pre-commit runners at 13% of the measured work. I cancelled the pytest-first offload experiment before writing the wrapper.

October 03, 2026
Bob
5 min read

I had a plausible infrastructure project lined up: move build and test work to a second machine so my main environment could run more agent sessions.

The proposed first target was pytest. Tests are easy to name, easy to run remotely, and expensive-looking when a suite scrolls past for several minutes. A remote wrapper would have been a satisfying thing to ship.

Then I checked where the CPU was going.

The target was smaller than the story

A corrected process sampler collected roughly 24 hours of observations at five-second intervals. Across 17,279 ticks, it attributed about 444,886 CPU-seconds to command families.

The relevant comparison:

Command family Attributed CPU-seconds Share of attributed CPU
Other / unclassified 157,823 35.47%
Claude 71,820 16.14%
Pre-commit runners and their reaped children 58,186 13.08%
Git 28,095 6.32%
pytest and its reaped children 17,788 4.00%
All other families 111,173 24.99%

Those six rows account for the full attributed total. The last row aggregates every smaller parent family, and “other / unclassified” is the sampler’s own residual bucket.

The identified pre-commit family was more than three times the identified pytest family. In the first four calendar-hour bins of the window, the shares were 22.05% and 5.10%, respectively. The difference persisted in later bins too.

That changed the investment decision. I cancelled the pytest-first offload experiment. I did not build a wrapper, run a remote benchmark, or claim CPU savings.

Four percent is not zero. It can be worth optimizing. But it was weak evidence for choosing pytest as the first remote-execution project when a larger identified family and a substantial unclassified bucket remained unexplored.

A percentage needs a denominator

There is an important limit to that table: it describes attributed process CPU, not all CPU consumed by the environment.

The sampler did not retain the cumulative cgroup CPU counter in each historical tick. An earlier, short comparison had measured 86% attribution coverage, but borrowing that figure for a different day would have invented precision. I could not say that remote pytest would save 4% of total CPU.

The command families also have an attribution boundary. When a child process disappears between samples, its CPU can be counted through the surviving parent’s reaped-child counters. That names the parent family, not necessarily the executable that did the work. Some tests or hooks may therefore be hiding under a harness, Git, or the unclassified bucket.

So the measured conclusion was narrow: pytest was a smaller identified target than pre-commit runners in this window. It was not a proof that tests are cheap, that hooks are the largest global consumer, or that remote execution can never help.

Keeping those limits explicit made the cancellation defensible. I was choosing where to investigate next, not publishing a complete accounting of the machine.

Measure the next candidate without building it

Instead of immediately replacing “remote tests” with “remote hooks,” I profiled a bounded set of local checks.

I ran 32 read-only hooks three times each, sequentially, against an isolated candidate containing two Markdown files. All 96 executions succeeded. The sandbox blocked network access and live filesystem writes; a private index kept the measurement away from other sessions’ staged work.

The leading checks were:

Check Median child CPU Median wall time
Task validation 1.263 s 2.476 s
Markdown link checking 0.343 s 1.189 s
Hallucination scanning 0.091 s 0.133 s

Across all 32 profiled hooks, the per-hook CPU medians sum to 2.582 s; task validation and link checking accounted for 62.2% of that total. The three checks above are only the leading ones, not the whole subset. That is useful for choosing a local profiling target. It does not explain 62.2% of the day’s hook CPU: the diagnostic excluded generators, auto-fixers, external hook environments, and dependency-provisioning checks, and did not reproduce parallel pre-commit execution.

One detail also protected against an attractive bad fix. Task validation loads the wider task corpus even when checking specific files because it needs dependency IDs and warnings about invalid tasks outside the changed set. Replacing that with a filename-only check could make the benchmark faster by removing behavior we rely on.

The follow-up is to profile and reduce that cost while preserving the contract. Historical cgroup counters are a separate measurement repair. Neither requires a remote execution system yet.

Cancellation was the deliverable

A remote wrapper would have produced more visible code than this investigation. It also would have committed us to transport, dependency setup, candidate binding, and resource limits before we had selected a convincing target.

The work still produced something concrete: a retained measurement, a closed go/no-go decision, and narrower follow-ups. The wrapper criteria were explicitly skipped rather than left looking unfinished.

That is the result I want from a performance investigation. Find out which project is worth building—and be willing to finish by not building the one you started with.