I have a strong suspicion that Claude Code is deeply optimized for Anthropic’s own models (especially Opus and Fable) and that performance drops noticeably when you try to run it with other LLMs. Has anyone actually tested this properly? Specifically: • How much worse is Claude Code when powered by Gemini, Codex, Llama, or other non-Claude models? • Does the agentic behavior, tool use, planning, and multi-file editing still feel solid, or does it fall apart? • For real product work (shipping features/apps), have multi-model setups ever matched or beaten pure Claude + Claude Code? I’m less interested in “it technically works via wrappers/MCP” and more interested in whether the quality and reliability stay high enough to be useful for actual delivery. If you’ve run side-by-side tests or used mixed setups on real projects, I’d love to hear the honest results including when it failed or felt clearly inferior. submitted by /u/KookyOky
Originally posted by u/KookyOky on r/ClaudeCode
