Original Reddit post

Google reports that Gemini Robotics ER 2 reaches 91.3% accuracy on moment finding, with a mean error under one second, but 57.4% accuracy when classifying overall task progress into five bands. These are different evaluations, so the numbers should not be compared as if they measured the same thing. Still, the gap reveals an important reliability problem. A robot may correctly notice the exact frame when coffee reaches the desired level while remaining uncertain whether the broader multi-step job is 40% or 80% complete. That matters for recovery, handoffs, billing, and deciding when a human needs to intervene. Should completion detection be a separate, independently tested safety layer rather than another output of the same generative planner? For physical agents, which is more important: precise local event detection, a calibrated estimate of global progress, or the ability to admit that the task state is ambiguous? submitted by /u/Crescitaly

Originally posted by u/Crescitaly on r/ArtificialInteligence