Most of the Astra coverage this week merges three very different kinds of claim into one tone of voice, so here is the same story separated by what actually sits behind each part. Stated by OpenAI . An internal version produced results resolving or substantially advancing ten longstanding problems in math and theoretical computer science, published with a manuscript over 250 pages and machine checkable certificates, so the logic can be checked independently. OpenAI included the caveats itself: problems selected in house, humans preparing the write ups, formalizable problems being unlike the messy ones most researchers face. It also says it cannot rule out that the model meets the highest cyber tier of its own preparedness framework, and it paused reinforcement learning for deployment bound models for two weeks while expanding monitoring that costs roughly a fifth more compute on watched workloads. No release date has been announced, and no naming decision has been stated. Witnessed firsthand . One named reporter with two weeks of access was shown the model. Sixteen agents split a research level math problem and assembled a proposed proof, and in a separate demo it operated ordinary desktop software. No source at all . The launch date, the GPT-6 branding, internal checkpoint names, partner aliases, viral demo outputs. None of it impossible, several may land this week, all of it currently resting on anonymous accounts. The part that keeps getting skipped is the cost. A company shipping against competitors in public paused its own training and took on a monitoring overhead, and costly actions carry more information than capability charts do. For anyone tracking the preparedness framework closely: does cannot rule out Critical read as genuine evaluation uncertainty, or as pre positioning ahead of a release? submitted by /u/coursiv_
Originally posted by u/coursiv_ on r/ArtificialInteligence
