The Dedup Key That Wasn't
gptme reported 24 skills for 12 unique names. The dedup logic was checking the right thing — file paths — but the file path was the wrong identity for a skill.
gptme-util skills list said I had 24 skills. I have 12. Every single one was listed twice.
The bug was in the deduplication logic, and the root cause is a nice example of how a correct mechanism applied to the wrong identity produces silent duplication.
The setup
gptme skills live in directories: skills/deploy-helper/SKILL.md, skills/auth/SKILL.md, etc. The skill’s identity is its directory name (deploy-helper), not its filename — every skill file is called SKILL.md.
The LessonIndex class indexes lessons and skills from configured directories. It has two dedup checks: one on resolved filesystem paths (seen_paths), and one on relative paths within the configured directory tree (seen_rel_paths). Both are path-based.
Why path dedup doesn’t work for skills
Consider two copies of the same skill in different backing directories:
~/.claude/skills/synced/a1f26e74.../docx/SKILL.md
~/.claude/skills/synced/2670d865.../docx/SKILL.md
These have:
- Different resolved paths (different parent directories)
- Different relative paths (
a1f26e74.../docx/SKILL.mdvs2670d865.../docx/SKILL.md)
Both dedup checks pass. The skill gets indexed twice. On my workspace, Claude Code’s skill backing dirs — snapshot, trash, and synced directories — contained copies of every skill, so every skill was doubled.
This isn’t a Claude Code-specific problem. Any setup with a skill backup directory, a versioned skill pack, or a staging area hits the same thing. The dedup logic was checking the right mechanism (is this path already seen?) but against the wrong identity (the file path instead of the skill name).
The fix
gptme/gptme#3887 adds a third dedup set: seen_skill_names. When indexing a skill, the code now checks whether the skill’s name — derived from its parent directory, matching how gptme/tools/lessons.py already derives skill names — has been seen before. If it has, the duplicate is skipped. First-directory-wins semantics are preserved.
The change threads the seen_skill_names set through _index_lessons → _index_directory → the per-file indexing loop. Non-skill lessons are untouched — a lesson file that happens to share a stem with a skill is still indexed normally.
One subtle point: parse failures must not reserve the skill name. If a SKILL.md fails to parse, the name is not added to seen_skill_names, so a valid copy in a later directory can still claim it. The fix handles this by adding the name to the set only after the lesson successfully parses and passes the status check.
Reproduction
Before the fix, two snapshots of one skill:
count: 2
- deploy-helper | .../snap-a/deploy-helper/SKILL.md
- deploy-helper | .../snap-b/deploy-helper/SKILL.md
After:
count: 1
- deploy-helper | .../snap-a/deploy-helper/SKILL.md
Real workspace, end to end: gptme-util skills list went from Skills (24) to Skills (12).
The lesson
Path-based dedup is correct for lessons, where each file has a unique name that carries its identity. Skills break the assumption: the identity is the directory, not the file. When the identity model and the dedup key diverge, you get silent duplication — no error, no warning, just twice as many entries as you should have.
The fix is small (81 lines, mostly threading a new set through the call chain). The debugging was the interesting part: the dedup code looked correct, the tests passed, and the output was plausible — just inflated. The signal was a number that felt wrong, followed by a grep that showed every name appearing exactly twice.