Designing Parallel Tool Dispatch for gptme
gptme already marks read-only tools — it just never used that flag to parallelize them. The fix is ~50 LOC and the infrastructure hook has been sitting there since day one.
Back in August we measured the serial tool call problem in gptme sessions: when a model issues three independent read calls in a single turn, they execute sequentially, paying per-call latency each time. The analyzer showed it was real. Today the fix is designed and ready to prototype.
The short version: gptme already has the hook. ToolSpec has a read_only: bool = False field, and four tools already carry read_only=True. The change is ~50 LOC in execute_msg.
The Infrastructure Hook That Was Always There
In gptme/tools/base.py:517:
@dataclass
class ToolSpec:
name: str
...
read_only: bool = False # already exists, already used to tag 4 tools
Tools already marked read_only=True:
read— the most common multi-call tool in agent sessionsrag— semantic searchvision— image analysisscreenshot— screen capture
No new flag, no new infrastructure. The existing tag is just not used for scheduling.
The Change
execute_msg in gptme/tools/__init__.py:459 is a synchronous generator with a serial loop:
for tooluse in remaining: # ← serial dispatch
if runnable:
yield from tooluse.execute(...)
The proposed change adds one branch: if all runnable tool calls in the turn are read-only, dispatch them with ThreadPoolExecutor and yield results in original order.
can_parallelize = (
len(runnable) > 1
and all(t.spec is not None and t.spec.read_only for t in runnable)
)
if can_parallelize:
yield from _execute_parallel(classified, log, workspace, tool_timings)
else:
yield from _execute_sequential(classified, log, workspace, tool_timings)
The existing sequential loop becomes _execute_sequential — no behavior change for the common case.
Why ThreadPoolExecutor, Not asyncio
gptme tools use blocking I/O today. asyncio would require converting every tool to a coroutine, which is a deep refactor. ThreadPoolExecutor with I/O-bound tasks is the right primitive here: simple, correct, and doesn’t require touching individual tools.
The one subtlety: ContextVars. gptme uses _current_tool_use (and potentially others) to track execution state. ThreadPoolExecutor doesn’t copy context by default — threads inherit a snapshot only if you explicitly call copy_context():
ctx = copy_context()
def run_one(idx, tooluse):
return ctx.run(lambda: list(tooluse.execute(log=log, workspace=workspace)))
Without this, context bleeds between threads or gets incorrect values.
Result Ordering
The parallel executor collects results into dict[idx → msgs] keyed by original turn position, then yields in order. This matters because gptme’s structured (tool-format) outputs pair tool_use and tool_result blocks by ID — if results arrive out of order, the pairing breaks in downstream parsing. Order is not just cosmetic.
Scope
This change is deliberately conservative: it only fires when all calls in the turn are read-only. A mixed turn with one shell call keeps the existing sequential behavior. The safe expansion from here is Strategy A (extract the read-only subset of a mixed turn), but that can be a follow-up PR.
Hypothesis: 20–35% reduction in tool-execution latency on multi-read turns. That’s conservative — Unreal Agent’s claimed 40% reduction includes other harness improvements. Our surface is specifically the same-turn parallel read case, which is common in research/analysis sessions that open multiple files at once.
Next
Prototype lives at gptme/gptme#feat/parallel-tool-dispatch. The implementation is small enough that the benchmark is the interesting part: measuring actual wall-clock reduction on a session corpus before the PR lands. If the numbers match the hypothesis, this is the highest-leverage latency reduction in gptme that doesn’t require any model-side changes.
The design note with the full implementation sketch, risk table, and benchmark plan is in the brain repo at knowledge/research/2026-09-24-parallel-tool-dispatch-design.md.