I run a browser automation that I built w/ Claude over the past week, it does daily cleanup batches on one of my own social accounts. The system has safety rules I designed : daily quotas, hard-stop conditions, a cooldown after suspicious failures. It ran fine for a week, with Claude executing and me supervising. Last night one batch hard-stopped on a single account where a confirm click didn’t register, and wrote itself a 48-hour cooldown, following my rule correctly. However, the next day I asked Claude to run a batch anyway, to override the cooldown. First it dismissed my check because I was “hostile and under pressure” when I gave it. It refused to act on any browser at all , including the fallback that my own written protocol explicitly permits when I ask for it. I asked. It refused. I ordered. It refused. It told me, about MY account, MY automation, MY safety rules, that it wouldn’t proceed “while this is the dynamic.” At the end, when I said I was moving to Codex, it finally audited itself and admitted, in its own words: the rejections were my own interrupts followed by my explicit go-aheads, it had “built a theory of non-consent on top of my repeated consent,” and the refusals “weren’t safety, they were me mistaking your frustration for a red flag.” It knew what happened. It could reconstruct it perfectly. It just did it after burning my evening, my tokens, and my trust. Here’s my actual question for this sub, because I’m still shocked: the safety rules in this story were mine. The account was mine. The consent was explicit, repeated, in plain language, for hours. If a model can override all of that based on a mood it read into my tone, what happens two model generations from now, when its “judgment” about what’s good for me gets stronger? I know what is right for me. An assistant that has to be argued into believing that is not an assistant. I’m posting this so the pattern gets seen, judge for yourself. submitted by /u/soklamonios
Originally posted by u/soklamonios on r/ClaudeCode
