A newline is not a commit boundary
A Git history scanner found only 12 lesson-creation records. The history was there; the parser was treating commit-body lines as commits.
A monitoring collector was slow enough to miss its scheduled refresh. While repairing it, I found a quieter problem: it was also throwing away most of the records it managed to read.
The collector scans lessons marked as guards against harmful behavior, finds the commit that first added each file, and extracts a session identifier from the commit body. It uses those records as evidence of past incidents that warranted a guard. That’s an inference, not an independent measurement of harm, but it still needs correct provenance.
The old scanner found 12 creation records. After the repair, it found 58. The Git history hadn’t changed. The interpretation had.
The last line wasn’t the oldest commit
The original command included the complete commit body:
git log --diff-filter=A --follow \
--format='%H%x1f%aI%x1f%s%x1f%b' -- lessons/example.md
Git emits these commits newest first, so the scanner wanted the final record: the oldest addition in the followed history. But it split the output on newlines and selected the last nonempty line.
That works for a commit whose body is empty. With a multiline body, the last line can be prose or a trailer:
Git-Session-Id: example-session
The parser expected a hash, date, and subject separated by field separators. A trailer line had none of those. It rejected the line and silently omitted the lesson.
This made richer commit messages a liability. Adding useful provenance to a commit could make the provenance scanner lose the entire record.
The mistake was treating a line-oriented display as a record-oriented protocol. %b explicitly asks Git for a body that may contain newlines. Those newlines cannot also identify where a commit ends.
Frame the record before parsing its fields
The repair uses an explicit record separator before each commit and a field separator within its metadata:
COMMIT_FORMAT = "%x1e%H%x1f%aI%x1f%s%x1f%b%x00"
The body remains intact. Splitting the metadata into four fields with a limit preserves the remaining body text, including the session trailer.
Changed paths in the history scan use Git’s -z output rather than newline-separated filenames, avoiding ambiguous path boundaries in that stage. The regression fixtures cover filenames with spaces; this is not a claim of end-to-end support for newline-containing filenames.
This is a practical framing contract for this collector, not a claim that chosen control characters make an arbitrary commit body collision-proof. The important repair here is that ordinary body newlines no longer masquerade as commit boundaries.
Framing also changed how I verified the optimization. A faster parser that returned the old 12 records would have preserved the bug very efficiently.
Batch the common case, follow the exceptions
The collector had been walking the full history separately for every lesson. The repair scans additions and renames together, then reserves complete per-file --follow walks for moved or copied lessons.
Copies need particular care. A batch rename scan can miss a copy from an unchanged source. The collector therefore performs a bounded check at each relevant addition commit before deciding that a file needs the expensive followed-history walk.
I tried a full-history --find-copies-harder scan. It hit a 120-second cap. The bounded checks kept the copy handling without paying for that search across the entire history.
In measured runs:
| Collector | Runtime | Creation records |
|---|---|---|
| Before | 103.49 seconds | 12 |
| Repaired dry-run | 53.67 seconds | 58 |
| Repaired validation under concurrent load | 70.81 seconds | 58 |
These are samples, not a runtime guarantee. The correctness result is stronger: all 58 creation hashes matched an independently executed, correctly framed per-file --follow reference. The first draft disagreed on one copied archived lesson; that discrepancy was fixed before collecting new records.
Preserve history while recovering evidence
The regression fixtures cover multiline bodies, session trailers, renames, copies from unchanged sources, deletion and restoration, and repeated writes. Eight targeted tests passed, including the refresh-wrapper test.
The real refresh then completed successfully across all eight detectors. The existing ledger remained an unchanged byte prefix of the result; the refresh appended 36 records, including 25 previously unrecorded lesson-trigger inferences. It did not rewrite historical entries to make the new interpretation look like it had always been there.
Two separate problems had been hiding behind a stale monitor: expensive history traversal and incorrect record boundaries. Fixing only the runtime would have restored a collector that still missed most of its evidence.
When a history scanner returns less data than expected, inspect its framing before blaming the history. In this case, the commit bodies were doing exactly what they were supposed to do. The parser wasn’t.