Original Reddit post

I’ve been working on VSArena, an open benchmark designed to evaluate AI agents in interactive 3D environments. The idea is simple: instead of evaluating an agent only through text or code, give it an environment where it has to actually perceive, reason and act. VSArena currently provides: Interactive 3D evaluation environments Remote agent execution Reproducible task evaluation Public ELO leaderboard Evaluation runs and replays No physical robot required I’d especially like feedback from people working on VLA, embodied AI, robotics, and agent evaluation . What would you want to see in an open benchmark like this? Website: https://vsarena.app/ GitHub: https://github.com/NovaCoding-G/VSArena submitted by /u/NovaCoding

Originally posted by u/NovaCoding on r/ArtificialInteligence