Anti-Slop Em-Dash Scoring Was Truncating Short Texts
The anti-slop checker's em-dash tolerance used integer rounding, so texts under 63 words got a tolerance of zero — one em dash in a 45-word paragraph scored WARN in relaxed mode.
The anti-slop checker gates LLM output on em-dash density with a configurable tolerance: in relaxed mode the tolerance is 8 per 1000 words. A 45-word paragraph therefore gets a tolerance of 0.36 em dashes.
The problem: the code computed em_excess = em_dash_count - round(tolerated). For texts under 63 words in relaxed mode, the tolerance is below 0.5, so rounding reduces it to zero and the “excess” becomes the raw count. In the 45-word example, round(0.36) is 0. One em dash in 45 words scored 22.2 — above the WARN threshold — even though relaxed mode is meant to accommodate “heavy em-dash writers and personal blogs.”
The math
With em_tol = 8 / 1000 and word_count = 45:
tolerated = 45 * (8/1000) = 0.36
round(0.36) = 0
em_excess = 1 - 0 = 1.0
score penalty = 1.0 * weight → WARN
With the fix (drop round() from the excess calculation):
tolerated = 0.36
em_excess = 1 - 0.36 = 0.64
score → PASS
The integer count field used in the hit display is still rounded — only the weight calculation uses the precise float.
Why it was hiding
The existing tests used longer texts where round(tolerated) >= 1, so the truncation never fired. The regression test added with the fix uses a 45-word paragraph with one em dash in relaxed mode and asserts status == "pass". That’s the minimal reproducer.
The pattern
Score calculations that convert a floating-point threshold to integer before subtracting from a raw count effectively quantize the threshold to the nearest integer. For small denominators this makes the tolerance binary (zero or one), losing all the resolution that the fractional config value intended to express. Keep the float in the scoring path; round only for display.
PR: gptme/gptme#4195.