Original Reddit post

I’ve been thinking about this a lot lately and wanted to get people’s opinions here. So basically, every major AI company now offers some form of “temporary” chat mode. OpenAI has Temporary Chat, Google Gemini has Temporary Chats, Anthropic has Incognito Chats. They all say the same thing: these chats won’t be saved to your history and won’t be used to train their models. Sounds great in theory, right? But here’s the thing that bugs me. If you actually read the fine print, “temporary” doesn’t really mean temporary. OpenAI is pretty upfront about it actually. Their own FAQ says that Temporary Chats are deleted from their systems after 30 days, and during that window they may be reviewed to monitor for abuse. So it’s only ephemeral from our perspective as users. The data still sits on their servers for a month. Google does something similar but keeps it for 72 hours instead. Anthropic’s Incognito chats aren’t used for training, but deleted conversations may still hang around in backups for 30 days before permanent deletion. And honestly? I have a hard time believing that companies spending hundreds of millions (or billions) on training the next generation of models are just throwing away all that conversational data. Like, that’s some of the most valuable, naturally occurring training data you could ask for. People typing real questions, real follow-ups, real corrections. You’re telling me they’re not at least skimming off some of that for RLHF or something similar? There was even a thread on r/OpenAI a while back where people noticed OpenAI was doing A/B testing on Temporary Chats, which makes you wonder what exactly they’re doing with that data if it’s not supposed to be used for anything. There’s also the whole legal angle. OpenAI was under a court order for a while (related to the NYT lawsuit) to retain consumer ChatGPT and API data indefinitely. They said they fought it and eventually got out from under that order, but it shows that “we delete your data” can be overridden by legal demands at any point. Now, where I do feel somewhat more confident is with API usage. OpenAI explicitly says API data is not used for training by default, and they back that up with enterprise privacy commitments. Anthropic says the same about their commercial products (Claude for Work, API, Claude Gov). The reason I buy this more is that these are B2B/Enterprise clients pushing massive volumes of sensitive company data through the API. If it came out that an AI provider was secretly training on API data, the enterprise contracts would evaporate overnight, and the lawsuits would be brutal. Plus, you’re literally paying for tokens, so there’s a cleaner commercial justification for not needing to squeeze training value out of it. But even there, some people have pointed out that the terms of service only restrict using data for “AI training” specifically, which doesn’t necessarily close the door on other uses like safety monitoring, abuse detection, or “product improvement” which can be pretty broadly defined. So I’m curious what people here actually think: Do you believe the claims that temporary chats are genuinely not used for training? Which companies do you trust more on this, and which ones don’t you trust at all? Is the API a different story in your mind, or do you think it’s the same situation, just wrapped in better marketing? Genuinely interested in hearing different perspectives! submitted by /u/EriksonThorsen

Originally posted by u/EriksonThorsen on r/ArtificialInteligence