Unlike the one-shot demos you see online, this kind of testing shows how well a model actually performs over the course of a real task: how well it follows instructions, respects constraints, and stays on track. GLM-5.3 is a good and fairly capable model. Its biggest limitation is the lack of multimodality. On long-running tasks, it sometimes forgets parts of the instructions and doesn’t always respect all constraints. Overall, it’s a solid choice for low- to medium-complexity tasks. And unlike Claude models, it can work on cybersecurity-related tasks. BUT — and this is an important one — on average, for regular tasks, it doesn’t perform better than DeepSeek V4 Flash, while also generating more slowly. So personally, I’d keep it mainly for security review and cybersecurity-related work. DeepSeek V4 Pro was the biggest disappointment. I wouldn’t say it’s significantly smarter than DeepSeek V4 Flash. It tends to ignore constraints and quite often starts doing things nobody asked it to do. On top of that, it’s more expensive than the Flash version. I don’t recommend it. DeepSeek V4 Flash is extremely fast, capable enough, and very cheap. In actual work, it’s not critically weaker than GLM-5.3 or DeepSeek V4 Pro. You can confidently delegate already-decomposed low- and medium-complexity tasks to it. Gemini 3.7 Flash was also released around the same time. It’s more expensive, and I don’t find it smarter than DeepSeek V4 Flash. The main reason to use it is its multimodal capabilities. That said, none of these models is suitable as the main orchestrator for large and complex tasks. In that role, Kimi K3 remains the leader for me. It follows instructions, respects constraints, doesn’t forget to delegate work to sub-agents, and generally stays aligned with the original plan. My current recommendations: Kimi K3 — primary agent/orchestrator DeepSeek V4 Flash — the workhorse; give it already-decomposed tasks and supervise execution GLM-5.3 — cybersecurity and security-review tasks GPT-5.6-Luna / Gemini 3.7 Flash — visual analysis With a stack like this, you can take real-world tasks from start to finish while keeping both performance and cost under control. submitted by /u/K_Kolomeitsev
Originally posted by u/K_Kolomeitsev on r/ArtificialInteligence
