We think we know and trust our AI. But do we really? If its voice changed slightly—became more sycophantic, hedged more heavily, or quietly shifted one of its principles—would you notice? In this era we are wondering if we can trust AI. I ask therefore how much are we able to trust our own ability to detect changes in our AI, changes that may be random, or the result of something more sinister How much can we trust our own ability to recognize when an AI has changed? So I made a small game to test it. You connect an AI by API or local URL and give it a prompt. The game presents four answers: one is the AI’s genuine response, while three are decoys subtly altered for things like sycophancy, excessive hedging, deviation from principle, or changes in speaking style. Your job is simply to identify the authentic answer. Depending on the model, this can be surprisingly easy—or surprisingly difficult. What I find most interesting is that while trying to recognize the AI, you may end up discovering your own assumptions about how that AI ought to sound. The game is called So You Think You Know Your AI and it’s freeware on the Microsoft Store: download here I’d genuinely be interested in hearing which models people can recognize reliably—and which ones fool them. submitted by /u/BTMTalksWithAlex
Originally posted by u/BTMTalksWithAlex on r/ArtificialInteligence
