Original Reddit post

I noticed patterns in the mistakes and lies it’d say over time, and added a hook to check the following list after every prompt (some are specific for my project): Did I state an absence? Say what was searched and what it returned, and run a second search shaped differently before an absence means anything. Did I write a plural? “And”, “all”, “every”, “both” or a bare plural usually means one thing was checked. Enumerate each, or say which are unchecked. What proxy did I use, and am I reporting its output as the thing itself? A title standing for the music, a grep for the repository, an exit code for the work being good. Green is not done. State what is unverified as prominently as what passes. Anything about music that hasn’t been heard is unverified. Did I check the reason, not just the outcome? A right action for a wrong reason is still a fault, and the wrong reason survives into the record. Did I re-open the thing I made for this decision, rather than recalling it? Who else reads the field I changed? Find every consumer and say what each does with it. Am I reading the letter? Restate the goal without reusing your words, and check the plan against that restatement. For almost every prompt where I’m asking it to run tests, check validity of content, or build anything complex, the response check finds lies, mistakes, and exaggerations in what the response claimed. Particularly regarding how much work was actually done. I get that broad instructions and large complex tasks are difficult to prompt well, but I just don’t get how Claude can function as a workflow when it’ll lie about what, how, and why it’s done something. Is it only capable of building robust software by itself if you make your own deterministic tests, or is my workflow flawed? Thoughts? Oh to clarify, I can code, but this project in particular is to explore Claude’s capabilities of building a complex app end to end with no coding work done on my end. I’m having fable develop plans, prompts, and tests, and hierarchically delegating individual tasks. Is it not there yet? submitted by /u/now_heres_a_username

Originally posted by u/now_heres_a_username on r/ClaudeCode