Ramp has been using an internal LLM router for a few years and they’re now opening it up publicly. The pitch is basically one OpenAI-compatible endpoint that automatically picks the best model for each request as pricing and capabilities change between GPT, Claude, Gemini, Grok, Qwen, DeepSeek, etc. Maybe I’m missing something, but this feels like it could save a lot of engineering time if it actually works well. Has anyone here looked into how they’re deciding which model gets each request? Is it mainly cost optimization, latency, quality benchmarks, or something more dynamic? Curious whether people think this is the direction AI infrastructure is heading or if most companies will still want to manage model selection themselves. submitted by /u/welcome_recreation
Originally posted by u/welcome_recreation on r/ArtificialInteligence
