Anyone building on a hosted model has had the moment where it feels different and nobody said anything. I wanted something more concrete than a feeling, so since Aug 20 I’ve been running the same private set of probes every day against 34 models from 15 labs, over their APIs, and comparing each model only with its own past. No rankings, no LLM judging the answers, just fixed checks graded by code. Every reading is hashed into a public transparency log the moment it’s taken, so nobody can argue later about when the baseline existed. The first real catch: on Sept 10, DeepSeek’s reasoner started using roughly 10 to 12x the thinking tokens it had used for the previous three weeks, on the same probes. I couldn’t find any announcement. That’s not the model getting worse, it’s the model getting slower and more expensive, which is the kind of change people describe as “it feels off” and then can’t prove. Nothing on the other 33 looks like a quiet drop in capability so far, and that’s a result too. Opus 5.5 has been on it since it launched. Code and methodology: https://github.com/piperoll/seismograph Live readings: https://seismo.piperoll.org/ Still early. Happy to answer questions on how it works or what it can’t detect yet. submitted by /u/Electrical_Rip892
Originally posted by u/Electrical_Rip892 on r/ArtificialInteligence
