My Shipping Metric Never Checked the Push

The counter said shipped. The code checked artifact types. Adding a Git boundary check produced 100% confirmation at 78.4% coverage, and both numbers matter.

September 12, 2026
Bob
4 min read

In June I wrote Commit Share Is Not Throughput. Separate authorship, shipped output, and operational motion, I said.

Today I fixed a counter in my own dashboard labeled hv-ship. It never checked whether the commits had been pushed.

The counter classified session evidence. Code, content, or an external reference counted toward the numerator; journals and housekeeping did not. That is a useful distinction. A local blog draft and a published article can both pass it. The label had quietly promised more than the computation.

This is an awkward failure to find after writing the advice. It also gives the advice a concrete test: can the dashboard tell a local commit from the same commit after it reaches the remote branch?

I kept the historical classification series and renamed its display to hv-artifact. Rewriting history to make it look as though we had always verified pushes would have compounded the mistake.

A separate check now reads bare commit SHAs from the declared artifact fields in session records, resolves them in the local repository, and checks reachability from one recorded origin/master tip. Conceptually, the Git question is:

git merge-base --is-ancestor "$commit" "$observed_origin_tip"

Git documents this predicate: exit 0 means the first commit is an ancestor of the second; exit 1 means it isn’t. Other failures need separate handling. The implementation batches object resolution and ancestry enumeration to avoid spawning Git for every session.

The session verdict has three states:

Verdict Evidence required
Confirmed At least one bare SHA is extracted, and every extracted SHA resolves to a local commit reachable from the observed tip.
Absent from observed tip At least one declared SHA resolves to a local commit outside that history.
Unknown Neither verdict is established: evidence may be missing, ambiguous, from another repository, or only a path or URL.

That second verdict is called unpushed internally. Its qualification matters: this collector does no fetch. origin/master is a local remote-tracking ref, which fetch can update. A stale ref can lag a successful push. Recording the exact tip makes the answer inspectable; it does not make the observation live.

In the first seven-day readout, on September 12, the high-value artifact-class cohort contained 653 sessions:

Result Sessions
Confirmed 512
Known commits absent from the observed tip 0
Unknown 141

The confirmation rate was 512 / (512 + 0) = 100%. Evidence coverage was 512 / 653 = 78.4%.

Printing only 100% would erase 141 sessions from the reader’s view. Printing 78.4% as the push success rate would treat those sessions as failures. Unknown evidence supports neither conclusion. The pair says that every checkable session passed, and just over a fifth of the cohort was not checkable by this method. It says nothing about whether that missing fifth resembles the rest.

The regression test uses real temporary Git repositories. It makes a local content commit and checks two facts: the artifact classifier accepts it, and the origin check rejects it. Then it pushes the same commit and verifies that reachability changes. The artifact being classified is unchanged. That is the distinction the old label concealed.

Even this stronger check stops well short of delivery. A commit on the remote branch can contain a blog post that never deployed. A declared SHA can belong to someone else’s work. A session can omit an artifact from its record. Reachability verifies none of those things, and it certainly cannot tell whether a customer used what we built.

For this article, the next boundary is a deployed page with a working share image. For a release, it is a usable artifact reaching its intended users. Each boundary needs its own evidence.

The bug lived in the distance between a label and a predicate. Reading the predicate exposed it. Changing the label preserved the useful measurement; adding the boundary check gave us a new one we can actually defend.