Why use FlexInference
By Aditya Perswal
Use FlexInference if you want a dead simple router that will just do what you ask, while making billing extremely clear and saving you money.
On tasks that can wait, either you pay about half, or you pay exactly what you would have paid anyway. Across everything running with a deadline now, requests pay 48% less and the first token arrives about 20% later. You tell us how long a request can wait to start, anywhere from five seconds to ten minutes. The integration is a base URL, a key, and one parameter.
Finding the 10x
Use it when you want to test out different models and go looking for the 10x or 100x savings. Somewhere in your workload there's a task that doesn't need Sonnet, and Kimi or DeepSeek would do it for a fraction of the price.
Swapping the model name is easy. The hard part is knowing the new model handles your workload without degradation or regressions. That's what we're building A/B testing for. It sends part of your traffic to the new model and scores it against the one you trust, on your own tasks.
Seeing what your models do
Use it when you want a better view into how your models work. We have observability into every model's latency and availability, connected to its provider, and into your own traces and logs when you want us to keep them. We store traces and request content only if you ask us to. When something is slow or expensive, the logs show which request it was and what it cost.
Billing you can read
Use it when you want to stop being upcharged random amounts on model pricing. Every model costs what the provider charges, passed through. Managed keys pay 10% on top-ups and nothing else. Bringing your own keys (BYOK) is entirely free.