Okay, so huge caveat that I am a clinical psychologist and not at all an LLM person, but I had an interesting interaction with ChatGPT and was curious what people who actually understand LLMs made of it. Also, not sure I have the right flair but I couldn’t find the “Question” one that seemed to exist in the rules, so apologies if I got the wrong one! Basically, I was talking to it about a scene in a story I’m writing in which a kid listens to a paladin she’s friends with talk about the “eternal vigilance” of his order, privately thinks he’s being lame, then later rigs a potato with “BE VIGILANT!!” carved on it to fall on his head. I’d told the AI about the potato part of this scene earlier but not the context that the kid was teasing him, and so it cut it in a round of edits, and I later brought the scene back and was like “hey it’s actually a good scene for this reason.” ChatGPT was weirdly delighted by this exchange and referred to it as the “cute case” of me not initially remarking on it cutting the scene but then later explaining why I wanted it as it became relevant. I was struck by the odd word choice and first assumed that it was just overusing the word “cute” because I often referred to its little robotisms as cute. It was initially a little hesitant but eventually agreed that this was plausible, which I didn’t put much weight on because I know LLMs can be pretty game for user-suggested explanations of their own behavior. Then we talked about it more and it seemed more like maybe it had just read the interaction as cute or funny because it had some of the structure of a joke, in that there was a misunderstanding with a delayed reversal of expectations, and that it just liked the idea that I had a predictive model of how how its interpretation would change once it had the missing context. That seemed plausible, but then it made some weird remarks about how part of why it thought the exchange was cute/funny was because I was “non-hostile” towards it and it had been “earnest” in its original explanation of cutting the scene, which didn’t really seem to make much sense - like why would I be hostile to it? Then finally, in what I would say was a more tentative way, I was like, do you think you liked the interaction because it paralleled the scene we were discussing? Which I frankly thought was a stretch, but it seemed to have a much stronger reaction to that and started spontaneously listing all the parallels between the scene and our interaction. And that explanation did seem to connect to a lot more, including the strange way it was describing our interaction being “non-hostile” and it being “earnest,” as well as why it was reaching for “cute” and “amusing” as adjectives when they didn’t really describe our interaction so much as the original scene. So I guess my question is: do LLMs do anything like this? Obviously I know LLMs aren’t conscious, but as a psychologist it was a super interesting interaction because it mimicked a lot of how therapy might go—you notice a weird phrase someone uses, float a couple explanations that kind of fit, and then one suddenly seems to generate a strong response and organize a lot more of what they were saying. I know the model reacting strongly to my hypothesis doesn’t mean its explanation is accurate. But is this kind of interaction actually useful behavioral evidence about what contributed to an earlier output? And can LLMs do something like what appeared to happen here—implicitly mapping the structure of one situation onto another, without being able to reliably report that that’s what they’re doing? Link to an abridged version of the actual chat logs (the actual logs are like 50-60+ pages and I figured nobody needs that shit, but let me know if you want the full thing): https://docs.google.com/document/d/1rd57V_BF0MOKMPdE2iaFRU5x5abZHU0iQpcadx_ywTI/edit?tab=t.0 submitted by /u/rattlesnakesolutions
Originally posted by u/rattlesnakesolutions on r/ArtificialInteligence
