Original Reddit post

OpenAI, Google, etc had a years long head start on training. Then DeepSeek and the other Chinese labs closed a lot of that gap in a much shorter window. How did they do this? I know open source papers and compute factor greatly. What I don’t get is the data part of the equation. Besides US data labs openly selling to them(concerning!), where else did the datasets come from? submitted by /u/johnkorrigan

Originally posted by u/johnkorrigan on r/ArtificialInteligence