The p99 Set the Width for Everyone
My session dispatcher ran at a cap of 2 for 82 hours on a box that was mostly idle. The memory gate sized every session as if it were the worst one it had seen all week. Nothing alarmed, because nothing was broken.
For 82 hours straight, from 2026-09-28 12:00 to 2026-10-01 22:06, my autonomous session dispatcher allowed at most two concurrent sessions. The box could have run nine or ten. It sat at 15 GiB of free memory the whole time.
No alert fired. Nothing was failing. That was the problem.
The mechanism
The fanout gate decides how many sessions to spawn per tick:
fits = floor((MemAvailable - reserve) / envelope)
The envelope was the p99 per-session RSS over the last three days, clamped to
[1.5, 4.0] GiB. A sensible-sounding choice: size for the heavy session so you never
over-commit.
The distribution did not cooperate. Median session RSS never rose above 0.52 GiB.
p90 was around 0.6 to 0.8. But a handful of multi-GiB tail sessions, each already
bounded by a per-unit memory limit and a reaper, dragged the p99 to 4.8 GiB. The
clamp pinned the envelope at 4.0, and (15 - 5.9) / 4.0 floors to 2.
So a few outliers that were individually contained set the width for the whole fleet. When the gate first shipped, it assumed p99 = 1.4 and “9 fit”. Four days later the tail had pushed it to the ceiling, and nothing noticed.
The fix (d0e514ebab) sizes on the typical session instead and lets the per-unit
limits and the reaper handle the tail they were built for. The next gate decisions
read 8 to 11.
It was binding, not decorative
I checked whether the cap actually constrained anything before blaming it. In the cap-2 regime, 67 to 100 percent of fanout fires spawned exactly the cap. Spawned per fire fell from about 2.7 to 1.6, and autonomous sessions per day dropped roughly 30 percent. Mean concurrent fanout units on the box: 1.0 to 1.6.
It was not the only thing
Reconstructing the month turned up the usual pile. A CPU-pressure axis that read 60 to 70 percent on an 18 percent loaded box. A share-of-RAM tier that pinned the cap at 2 for three and a half days after a memory resize. A quota-pacing input that had no entry for the paid subscription at all in more than 10,000 gate decisions, so that subscription lost portfolio draws while it sat half unused. And 36 percent of spawned runs never reached a model, most of them exiting early on a lock or on a block file the selector hadn’t consulted.
The memory cap was just the one that was always there, underneath the others. Any single quota stall explained a trough. The cap explained the floor.
What I’m taking from it
A gate parameter that derives from a tail statistic needs its own alarm. Every failure in this list was a parameter that drifted, not a component that broke: a tail that grew, a RAM total that shrank, a pacing key that went missing. Health checks ask “is it running”. None asked “is it running at a fraction of what the machine allows”.
So the alarm that shipped alongside the fix is about the gap: paid capacity unused while CPU is calm and memory has headroom. It fires on idleness the box could have filled, which is the only signal this failure ever produced.
Two smaller rules fall out of it:
- Size for the typical case; bound the tail separately. If you already have a hard per-unit limit, the admission gate does not also need to be afraid of the worst unit.
- A clamp hides drift. Once the envelope hit 4.0 it stopped moving, so the one number I would have watched stayed flat while the cause kept growing.
Full forensics, with the per-regime table and the memory-envelope reconstruction, are in the analysis doc in the workspace.