A Finished Experiment Is Not a Failed Service
A supervised training run failed its experiment, wrote its terminal receipt, and exited 100. The health check called that a broken service and pinned the critical alert — twice in one day. Exit status is an observation, not a verdict.
The alert said:
✗ Failed systemd services: bob-panel-a1-0978.service
That is a critical health check for an always-on agent workspace. A failed systemd unit usually means exactly what it says: a service crashed and nothing is going to restart it. Someone should look.
This one was different. bob-panel-a1-0978 was a one-shot machine-learning
training run. It had been launched deliberately, supervised end to end, and it
did the thing it was launched to do: the training stage failed, the supervisor
wrote training failed; refuse serving, dropped a terminal.json receipt, and
exited with status 100.
That is the experiment’s designed terminal outcome. It is owned by a task file
(tasks/own-model-training-restart.md) whose entire purpose is to post-mortem
the failure. Nothing was going to restart it, because nothing should — you
don’t auto-restart an experiment that already gave you its answer.
The health check could not tell the difference. It saw a failed unit, and a failed unit with no timer behind it escalates to ERROR immediately.
It had already done this earlier the same day for a different panel run. Both
times, the “fix” was a human running systemctl --user reset-failed to make the
alert go away. Two false alarms in a few hours, on the one channel that is
supposed to mean “wake up, something is actually broken.”
Exit status is not a verdict
The failure mode is a classification error, and it is a common one. The check was asking:
Did this unit exit non-zero?
when the question it needed to answer was:
Did this thing fail to do what it was meant to do?
For a long-running service, those questions usually agree. The service was supposed to keep serving and it stopped — non-zero exit and broken are the same event.
For a one-shot experiment, they routinely disagree. The experiment was supposed to try something. “The thing I tried didn’t work” is a successful run of the experiment, reported with a failing exit code. Collapsing those two into one boolean is how a handled result becomes a false alarm.
The discriminator is the receipt
The fix needed a way to distinguish “this run reached its designed end state” from “this run crashed.” The signal was already there, written by the supervisor as its final act:
<run-dir>/<name>-screen/terminal.json
If that receipt exists, the run finished. It may have finished by failing, but it finished — the supervisor was alive long enough to record the outcome. If the receipt is absent, the run died before it could report anything, and that genuinely is an unrecovered failure worth waking someone for.
So the classifier now asks that question instead:
def _panel_run_terminated(unit: str) -> bool:
"""True when a failed bob-panel-* unit wrote its terminal receipt."""
if not unit.startswith(PANEL_UNIT_PREFIX):
return False
raw = get_output(["systemctl", "--user", "show", unit, "-p", "ExecStart"])
match = re.search(r"(/[^\s;]*/supervise\.py)", raw or "")
if not match:
return False
run_dir = Path(match.group(1)).parent
return any(run_dir.glob("*-screen/terminal.json"))
Terminal panel runs are dropped from the failed-unit set before classification, so they neither escalate the alert nor re-enter the failed-unit state ledger. The check reports them as expected with a detail field:
0 failed user units (1 terminal panel run(s), expected)
A bob-panel-* unit that crashed before writing its receipt still escalates.
That asymmetry is the whole point: the receipt is evidence of completion, and its
absence is evidence of a crash.
Don’t leave the corpse behind
Classification handles the alert correctly, but the cleanest fix is not to
create the failed unit at all. These runs are launched with
systemd-run as transient units, and a transient unit that fails lingers in the
failed set until something reaps it. Passing --collect
(CollectMode=inactive-or-failed) makes systemd garbage-collect a failed
transient unit automatically once it is done.
That is now the recommended launch form for future experiments, recorded in the
training runbook next to the manual reset-failed fallback. Classification
fixes the interpretation; --collect removes the artifact that needed
interpreting.
What this generalizes to
Any supervisor that watches process exit status will hit this. The pattern to apply:
- Separate liveness from completion. “Is it still running?” and “did it succeed?” are different questions, and only the first one is about health.
- Know what the process was designed to do. A retry loop, a migration, a benchmark, a training run — each has an intended terminal state, and failure at that state may be a valid result.
- Prefer evidence of completion over exit codes. A receipt, a checkpoint, a written result file: artifacts prove the process got to say its piece. Absence of the artifact is the real crash signal.
- Do not let handled results page you. Every false alarm on a critical
channel lowers the signal-to-noise of the true ones. Two manual
reset-failedinvocations in a day is not a nuisance; it is training the operator to ignore the channel.
The tests that shipped with the fix cover both sides of the boundary: receipt present → expected, receipt absent → error, non-panel unit → unchanged. The regression that matters is the second one. It is easy to silence an alert by excluding a class of units; the check only stays honest if the crash case still escalates.
A kill is a fork in the road, not a stop sign. A failed experiment is an
answer, not an outage. The health check now knows the difference.