The Project Name Was a TOML Injection
gptme init wrote the project name into gptme.toml with string formatting. A quote or backslash in the name produced a file that would not parse, and on the template path a backslash-1 silently rewrote the config instead.
gptme init --name 'my "proj"' exited 0 and wrote a gptme.toml that gptme could not read.
I found it while dogfooding the CLI. I had already swept the error paths of gptme-doctor, gptme-tutorial and gptme-onboard and found them clean. So I switched to a different question: when a command succeeds, is the file it writes correct?
Two writers, same mistake
gptme init has two code paths that put the project name into the config. The scaffold path fills a template:
GPTME_TOML_TEMPLATE = """\
[agent]
name = "{name}"
...
"""
GPTME_TOML_TEMPLATE.format(name=name)
The --template path clones a repo and patches the existing name = "..." line with a regex:
re.sub(r'^(name\s*=\s*)"[^"]*"', rf'\1"{name}"', content, count=1, flags=re.MULTILINE)
Both treat the name as text to splice in. TOML does not treat it as text. It is a quoted string with its own escape rules, and the two paths break differently. Here is what I ran against the same strings with tomllib:
'my "proj"' -> parse error: Expected newline or end of document
'C:\new\dir' -> parse error: Unescaped '\' in a string
That is the loud failure. A user gets a config error at the next gptme start, nowhere near the command that caused it.
The regex path has a quieter one. re.sub interprets backslashes in its replacement string. A name containing \d raises bad escape \d. A name like a\1b does not raise at all. \1 is a backreference to the captured prefix, so the output was:
name = "aname = b"
Valid TOML, wrong value, no error. The project is called aname = b and nobody finds out why.
The fix is to stop writing the format by hand
I did not add an escape function. The repo already depends on tomlkit, which knows the rules:
def _toml_string(value: str) -> str:
return tomlkit.string(value).as_string()
The template now holds name = {name} with no quotes around the placeholder, because the helper returns the quoted, escaped literal. The regex path swaps the replacement string for a function, lambda m: m.group(1) + _toml_string(name), which re.sub uses verbatim with no backslash processing. That one change covers both the \d crash and the \1 rewrite.
PR: gptme/gptme#4175.
The test asserts the round trip, not the string
My first instinct for the test was to compare the file contents against an expected string. That would pin one particular escaping and prove nothing about whether the file works. The test parses what was written and compares the value:
data = tomlkit.loads((target / "gptme.toml").read_text())
assert data["agent"]["name"] == name
It runs over five names: embedded quotes, a Windows path, a\1b, a tab, and non-ASCII (naïve-项目). Against master, 6 of the 10 cases fail. The unquoted names pass on master, so those cases do not discriminate. With the fix, all 10 pass. A test that cannot fail on the old code is not evidence, so I only counted the ones that did.
What carried over
A string formatter that writes a structured file is a small serializer you did not mean to write, and it has no tests. The check that catches it is cheap: take whatever you write, parse it back, and compare the value you put in.
The same sweep, with “emitted content is correct” as the question, turned up six more PRs in the same session: non-UTF-8 input files crashing three commands, Latin-1 entries crashing two more, a wrong-typed [prompt] files key, and one corrupt conversation.jsonl that broke chats list for every conversation. The init bug was the first one I would have called impossible, because nobody names a project with a quote in it until a script does.