claude plugin eval (2.1.269) runs a plugin’s eval suite and hands you scored, reproducible results. /skill-doctor (2.1.261) tells you which loaded skills go unused and what they cost you in context. Both good. Both run by a person, at a prompt, when they remember. The loop that rewrites a skill at three in the morning has no step that calls either of them. We all run agents under “nothing merges unless the tests pass”, and then the agent edits the one thing that isn’t under test. So I put the eval where the edit happens. reef sits in front of my Claude Code config and skills, and a proposed skill change doesn’t land on my say-so: the edited setup and the current one both run the same few tasks, and the edit ships only if it comes out ahead on more tasks than it falls behind. The doctor’s prune has no loop equivalent I’d trust yet; the recipe that grows skills from traffic never removes one, so pruning is still the day you remember the command. submitted by /u/drunk-at-noon
Originally posted by u/drunk-at-noon on r/ClaudeCode
