I work in regulatory research/compliance, mostly around chemicals, GHS, product compliance and regulatory intelligence across the Americas, with some broader global work. Over the last couple of years I’ve been experimenting a lot with LLMs, RAG, automation, structured regulatory data, prompt engineering, etc. And I keep running into the same contradiction: The work is almost absurdly well suited for AI — huge volumes of amendments, cross-references, transition periods, substance lists, definitions, jurisdiction-specific requirements, historical versions — but it’s also exactly the kind of work where a plausible-sounding 5% error rate is completely unacceptable. What interests me isn’t really “Can ChatGPT summarize a regulation?” Obviously it can. I’m more interested in whether anyone has built a workflow where the model can reliably distinguish between things like: what the law actually requires; what an authority merely recommends; what changed versus the previous version; whether an amendment modifies a list, a classification methodology, or only administrative language; whether two apparently related regulatory instruments actually operate together; and, crucially, when the model should simply say: I don’t have enough evidence to conclude this. My suspicion is that the winning architecture for regulatory AI won’t be a gigantic chatbot that “knows the law.” It’ll be something much more constrained: retrieval + structured regulatory data + deterministic rules + LLM reasoning only where ambiguity genuinely exists. Curious whether anyone working in RegTech, legal AI, regulatory intelligence or compliance has reached the same conclusion. What are you actually trusting LLMs to do today — and what do you absolutely refuse to delegate to them? submitted by /u/bluephoenix137
Originally posted by u/bluephoenix137 on r/ArtificialInteligence
