Original Reddit post

I run a small event staffing/activation agency as a solo operator, and Claude has basically become my back office: email, vendor communication, applicant database, calendar, etc. It’s a huge reason I’m able to operate at this size, but the confident mistakes are becoming a serious problem. A couple have nearly cost me deals. Some examples: • Told me a vendor had gone silent when she had answered every question that morning. Claude read an email search preview instead of the full thread and treated it as complete. • Told me an email was “staged,” so I went looking in Gmail for a draft that had never actually been created. • Referred to my suppliers as “brand partner candidates,” which completely changed the business context. A vendor gets paid by me; a sponsor pays me. I spent two days planning around an opportunity that didn’t exist. • Told me a table in my database didn’t exist without actually checking. It was there. The pattern I keep seeing is: Claude states things as facts that it did not actually verify. It seems to get worse during long sessions, where earlier conversation/context starts being treated like a source of truth. I’ve built some guardrails around it: a master state file for business context between sessions, fact labels like CONFIRMED / STATED / ASSUMED / OPEN, and rules like “never treat an email preview as the full thread” and “never claim an action happened without tool output confirming it.” That helped a lot. I went from several errors a week to a few a month, but I’m trying to figure out how people are making this reliable enough for actual business operations. For anyone using Claude this way: What are you using for durable memory/state across sessions? Has anyone built a verifier/checker that validates claims against tool output before Claude gives you an answer? Is there a reliable way to force Claude to actually retrieve/check the source instead of reasoning from previews or old context? What other guardrails or architecture changes have made a meaningful difference? I’m especially interested in hearing from people dealing with invoices, vendors, contracts, clients, deadlines, email, calendars, databases, etc. Situations where there’s a real consequence when the AI confidently gets something wrong. If you’ve dealt with this and found a setup that actually works, I want to hear how you built it. submitted by /u/JacktooGroovy

Originally posted by u/JacktooGroovy on r/ClaudeCode