Original Reddit post

I’ve been digging into AI training/privacy recently and some of the numbers are pretty wild. The UK’s ICO says generative AI training involves “vast amounts of personal data”, often processed without people knowing it’s happening. It says web-scraped training datasets can contain information relating to millions, if not billions, of people. And “public” doesn’t necessarily mean harmless. Research has demonstrated neural networks memorising unique information like names and IDs even when it appeared in just ONE training sample. The ICO gives a good example: someone posting about a doctor’s visit in 2020 probably wasn’t expecting that post to be scraped years later to train an AI model. Obviously this doesn’t mean ChatGPT has memorised everyone’s private information — newer research actually suggests some claims around PII memorisation have been overstated. But it made me wonder: how many people actually know they can object/opt out with some AI companies? The problem is every company has a different process, form or email address, and policies change. So I’ve built Don’t Train Me to automate the process and periodically resubmit opt-out requests: https://donttrainme.com/ I’m still very early with it, so genuinely interested in feedback — particularly whether people here actually care about opting out of training, or whether you consider public internet data fair game for AI. submitted by /u/Thin_Rush8229

Originally posted by u/Thin_Rush8229 on r/ArtificialInteligence