Original Reddit post

An AI system can make a workflow faster while making its mistakes harder to catch. Both things can be true at once. Picture an agent updating internal records, a thousand changes in the time a person used to need for ten. Each one is logged and can be rolled back. Sounds manageable, until other systems start acting on those records in real time, sending invoices, changing permissions, notifying customers. Rolling back the original record doesn’t necessarily undo what it already triggered elsewhere. Beyond whether it completed the task, how far can an error travel before someone detects and contains it? I’d want agentic workflow evaluations to measure three intervals, alongside accuracy and task time: how long until a consequential error gets detected, how long until further action stops once it’s caught, and how long until the downstream effects, including in connected systems, actually get repaired. Two systems can share the exact same error rate and still carry completely different risk. One stages changes for review before anything fires. The other executes immediately across several services. Same accuracy on paper. Very different exposure in practice. This also complicates what “reversible” even means. A database entry can revert cleanly while its consequences don’t. Fixing an invoice doesn’t un-send the email someone already read. Restoring access doesn’t give back the hours someone lost locked out of their account. I get the objection to adding checkpoints everywhere: delays cost money, frustrate users and create their own harm. Require approval for every minor action and you risk turning review into rubber-stamping. So I’d put the safeguards where consequences actually expand: external communication, money transfers, access changes, anything that triggers another system downstream. Limited permissions, staged execution, transaction caps, a pause before actions reach external systems, depending on the workflow. Speed can still make things safer, if the time saved buys people room to actually investigate. But if every bit of that savings just gets converted into more throughput, that room disappears. What I actually want to see measured: how much consequential activity a system can generate in the window between an error happening and someone containing it. If you’re deploying agents, are your evals measuring that window and what it touches downstream? Where does this framing break down, or miss something that actually matters? submitted by /u/OkyEscritora

Originally posted by u/OkyEscritora on r/ArtificialInteligence