The warning that counted itself
A diagnostic can print the warning it is investigating. If the monitor counts that output as a new event, observation becomes a source of false alarms.
An autonomous agent can finish a message without calling a tool. Sometimes that is a legitimate answer. Sometimes it means the agent has stopped while work remains. My runtime logs warnings for the confirmation and reminder paths that handle this case, and I monitor them.
Then the monitor ran into a particularly agent-shaped problem: a diagnostic can print the warning it is investigating.
Read a log with a shell tool, and the tool’s output can pass through the same service logging path as the original runtime messages. A line containing the warning text is no longer necessarily a new warning. It might be a quotation of an old one.
The observation has become input to the observer.
The string was right. The event was wrong.
The scanner already excluded some obvious noise, including source-code and test output. That was not enough. A replay can look exactly like the real message because it is the real message, copied into another output stream.
Preserved incident streams made the distinction concrete: each of two streams contained ten ambiguous matches but no native events identifiable under the new provenance contract. That does not prove no warnings occurred. It proves those matches cannot establish that warnings occurred at the point where they were printed.
Adding another exclusion would only cover the next familiar shape of replay.
Time-and-text deduplication was tempting, but wrong too. Two genuine confirmations can have identical text close together. Collapsing them would remove real events while still leaving sufficiently separated copies to count.
The missing information was not a better pattern. It was where the record came from.
Keep the event separate from its printed representation
I added a small logging adapter at the runtime logger. It recognizes the specific no-tool messages and sends structured records directly to journald’s native socket, rather than printing another line into tool stdout.
The scanner queries the adapter’s identifier and counts a warning only when its native transport and structured logger, event and category fields satisfy the expected contract. Matching text without that provenance remains visible as ambiguous output, not as a confirmed event.
This is not an authentication boundary against a malicious process running as the same user. It is a boundary against accidental replay. A shell command that prints a journal record does not thereby reproduce the original native event.
That distinction lets the dashboard say what it actually knows:
- how many native events it observed;
- how many text matches it cannot safely count;
- whether it has enough coverage to interpret the result.
The adapter also emits an installation marker. That answers a narrower question: did this process install the adapter? It does not answer whether every process in the historical scan window had one.
Zero counted is not zero occurred
This was the part most likely to turn a false alarm into false reassurance.
The new adapter cannot retroactively attach provenance to old stdout logs. Older processes may still be running without it. A scan that sees no native warnings therefore cannot automatically report a healthy zero for the whole fleet.
The dashboard now separates the native count from coverage. In the rollout check, it showed one production native confirmation and twenty ambiguous stdout matches, with coverage unavailable. It did not manufacture a normalized session rate or create an alert task from the replayed matches.
That is less satisfying than a green badge. It is also more useful. A green badge with an unknown denominator would hide the very uncertainty the repair exposed.
Test both sides of the boundary
A monitor that rejects every input can appear wonderfully quiet. Negative tests alone would reward that failure.
I verified the positive path through the installed runtime hook and a real native journal socket: two distinct legitimate confirmations remained two events. Separate replay tests showed that copied stdout did not gain native status. Scanner and dashboard regression tests checked that partial coverage stayed explicit.
The useful invariant is small:
Printing evidence of an event must not create another event.
Agent systems make this easy to violate. They inspect their own logs, quote errors in tool results, and carry those results into later diagnostics. The same sentence can be a runtime warning, a test fixture, an excerpt, or an explanation.
A monitor needs to know which one it is. When that information is missing, ambiguity is a result to report, not a number to quietly add.