# Why use FlexInference

Published 2026-08-11.

Use FlexInference if you want a dead simple router that will just do what you ask, while making billing extremely clear and saving you money.

On tasks that can wait, either you pay about half, or you pay exactly what you would have paid anyway. Across everything running with a deadline now, requests pay 48% less and the first token arrives about 20% later. You tell us how long a request can wait to start, anywhere from five seconds to ten minutes. The integration is a base URL, a key, and one parameter.

*Two bars for the same model. claude-fable-5 costs ten dollars per million input tokens directly, and five dollars per million with a start_within deadline.*

## Finding the 10x

Use it when you want to test out different models and go looking for the 10x or 100x savings. Somewhere in your workload there's a task that doesn't need Sonnet, and Kimi or DeepSeek would do it for a fraction of the price.

Swapping the model name is easy. The hard part is knowing the new model handles your workload without degradation or regressions. That's what we're building A/B testing for. It sends part of your traffic to the new model and scores it against the one you trust, on your own tasks.

## Seeing what your models do

Use it when you want a better view into how your models work. We have observability into every model's latency and availability, connected to its provider, and into your own traces and logs when you want us to keep them. We store traces and request content only if you ask us to. When something is slow or expensive, the logs show which request it was and what it cost.

## Billing you can read

Use it when you want to stop being upcharged random amounts on model pricing. Every model costs what the provider charges, passed through. Managed keys pay 10% on top-ups and nothing else. Bringing your own keys (BYOK) is entirely free.
