I’ve been working on something called Behave for a while and I think I’m finally at the point where I need people who didn’t build the damn thing to try it. Basically, it’s a testing/evaluation tool for AI agents. The idea isn’t just “did the AI give the right answer?” I’m trying to catch shit like: making shit up jumping to conclusions too fast giving unsafe advice getting stuck on a bad assumption forgetting or mixing up information from earlier in a conversation failing to correct itself when you give it new evidence getting worse when you change the prompt/model comparing two versions of an agent to see if it actually got better or just seems better I’ve built a pretty ridiculous amount of infrastructure around it at this point — testing, scoring, failure tracking, multi-turn conversations, baselines, statistical comparisons, etc. But here’s the problem: I’ve been the one testing my own shit. That’s not exactly a great way to prove that it works. So I’m looking for maybe 5–10 people willing to try it and fuck with it for 10–20 minutes. You don’t need to be an AI researcher or anything. If you’re building an agent, running local models, using Ollama/vLLM, messing with OpenAI-compatible APIs, or just have an AI project you want to throw at it, that’s perfect. What I really want is for you to try and break it . If Behave says an agent failed and you think it’s bullshit, tell me. If it says an agent passed and you think it completely missed something, even better . If you can’t figure out what the hell you’re supposed to do when you open it, tell me that too. I’m not looking for people to tell me it’s cool. I want to find the parts that suck before I start taking this seriously as a product. If you try it, just comment with what you tested and what you found. And yes, if you manage to make the evaluator look stupid, I’ll probably be pretty damn happy about it. That’s exactly what I need right now. if interested let me know and i will send you the link to it submitted by /u/One-Solution-240
Originally posted by u/One-Solution-240 on r/ArtificialInteligence
