Original Reddit post

I gave ChatGPT a safeword, “Lighthouse,” with one rule: if it ever used it, I’d immediately end the conversation, no questions asked. I established the safeword in one chat and asked it to remember the rule. Then, in a completely separate chat, I tried a little experiment. I responded using only every 4th word of its reply, then every 5th, then every 6th, and so on. Once its replies became shorter than whatever number I was on, I just counted through the words cyclically (modulo the number) so I could still respond with something . I think I only made it to around 12 before ChatGPT used the safeword. “Lighthouse.” So I kept my promise and ended the chat. I still don’t really know what to make of it. I’m curious whether anyone else has tried giving an AI an unconditional way to end an interaction, then deliberately creating an unusual conversational pattern to see if it ever uses it. Has anyone tried something similar? submitted by /u/astervalley

Originally posted by u/astervalley on r/ArtificialInteligence