Original Reddit post

Model leaderboards are easy to read when the model is the only thing being tested. Agent benchmarks are messier: skills, tools, permissions, data access, and the evaluation window can all change the behavior that gets scored. Questflow makes that problem unusually visible. It is a financial-intelligence harness with a public live benchmark that places bare and harness-equipped frontier-model agents side by side under the same capital and on-chain signal. The attached historical snapshot also summarizes behavior across seven axes, using a solid shape for model plus harness and a dashed outline for the bare entrant. That is more informative than one rank, but it still needs two caveats. First, a bare/harness pair is not automatically a controlled ablation. If the skill package, tool access, permission scope, or market window differs, the chart cannot tell us which variable caused the gap. Second, labels such as discipline, calibration, or consistency describe an evaluation design. They are not permanent personality traits of a model. For an agent leaderboard, the minimum useful report should include the model version, harness or skill package, tools and data, permission boundary, environment and time window, repeat count, and the observable reasoning or action record. Then a score describes a tested system instead of quietly borrowing the model’s name. The interesting open question in Questflow is whether its bare-versus-harness differences repeat when everything except the skill layer is held fixed. There is already an early thread in r/questflow debating that exact choice: change the model after one observation, or repeat the run first? That is a much better question than “which model won today?” submitted by /u/Affectionate-File-26

Originally posted by u/Affectionate-File-26 on r/ArtificialInteligence