Pricing

See how much each one of our models cost across different providers.

An indented row under a model is its rate once a request's input passes the size shown.

The two cache write columns are the provider's 5-minute and 1-hour cache rates. A dash means the provider publishes no rate for that line on that route.

OpenAI

OpenAI
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gpt-3.5-turbo$0.50$1.50———
gpt-4.1$2.00$8.00$0.50——
gpt-4.1-mini$0.40$1.60$0.10——
gpt-4.1-nano$0.10$0.40$0.025——
gpt-4o$2.50$10.00$1.25——
gpt-4o-mini$0.15$0.60$0.075——
gpt-5$1.25$10.00$0.125——
gpt-5-mini$0.25$2.00$0.025——
gpt-5-nano$0.05$0.40$0.005——
gpt-5-pro$15.00$120.00———
gpt-5.1$1.25$10.00$0.125——
gpt-5.2$1.75$14.00$0.175——
gpt-5.2-pro$21.00$168.00———
gpt-5.4$2.50$15.00$0.25——
Above 272K input tokens$5.00$22.50$0.50——
gpt-5.4-mini$0.75$4.50$0.075——
gpt-5.4-nano$0.20$1.25$0.02——
gpt-5.4-pro$30.00$180.00———
Above 272K input tokens$60.00$270.00———
gpt-5.5$5.00$30.00$0.50——
Above 272K input tokens$10.00$45.00$1.00——
gpt-5.5-pro$30.00$180.00———
Above 272K input tokens$60.00$270.00———
gpt-5.6$4.00$20.00$0.40$5.00—
Above 272K input tokens$8.00$30.00$0.80$10.00—
gpt-5.6-luna$0.20$1.20$0.02$0.25—
Above 272K input tokens$0.40$1.80$0.04$0.50—
gpt-5.6-sol$4.00$20.00$0.40$5.00—
Above 272K input tokens$8.00$30.00$0.80$10.00—
gpt-5.6-terra$2.00$12.00$0.20$2.50—
Above 272K input tokens$4.00$18.00$0.40$5.00—
gpt-6-astra$10.00$50.00$1.00$12.50—
Above 272K input tokens$20.00$75.00$2.00$25.00—
gpt-6-luna$0.10$0.50$0.01$0.125—
Above 272K input tokens$0.20$0.75$0.02$0.25—
gpt-6-sol$2.00$10.00$0.20$2.50—
Above 272K input tokens$4.00$15.00$0.40$5.00—
o1$15.00$60.00$7.50——
o3$2.00$8.00$0.50——
o3-mini$1.10$4.40$0.55——
o3-pro$20.00$80.00———
o4-mini$1.10$4.40$0.275——

Anthropic

Anthropic
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
claude-fable-5$10.00$50.00$1.00$12.50$20.00
claude-fable-5-1$10.00$50.00$0.25$12.50$20.00
claude-haiku-4-5$1.00$5.00$0.10$1.25$2.00
claude-opus-4-5$5.00$25.00$0.50$6.25$10.00
claude-opus-4-6$5.00$25.00$0.50$6.25$10.00
claude-opus-4-7$5.00$25.00$0.50$6.25$10.00
claude-opus-4-8$5.00$25.00$0.50$6.25$10.00
claude-opus-5$5.00$25.00$0.50$6.25$10.00
claude-opus-5-5$4.00$20.00$0.20$5.00$8.00
claude-sonnet-4-5$3.00$15.00$0.30$3.75$6.00
Above 200K input tokens$6.00$22.50$0.60$7.50$12.00
claude-sonnet-4-6$3.00$15.00$0.30$3.75$6.00
claude-sonnet-5$2.00$10.00$0.20$2.50$4.00

Google

Google
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gemini-2.5-flash$0.30$2.50$0.03——
gemini-2.5-flash-lite$0.10$0.40$0.01——
gemini-2.5-pro$1.25$10.00$0.125——
Above 200K input tokens$2.50$15.00$0.25——
gemini-3-flash-preview$0.50$3.00$0.05——
gemini-3.1-flash-lite$0.25$1.50$0.025——
gemini-3.1-pro-preview$2.00$12.00$0.20——
Above 200K input tokens$4.00$18.00$0.40——
gemini-3.5-flash$1.50$9.00$0.15——
gemini-3.5-flash-lite$0.30$2.50$0.03——
gemini-3.6-flash$0.75$3.75$0.075——
gemini-3.7-flash$0.75$3.75$0.075——
gemini-3.8-flash$0.75$3.75$0.075——

Amazon Bedrock

Amazon Bedrock
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
claude-fable-5$10.00$50.00$1.00$12.50—
claude-haiku-4-5$1.00$5.00$0.10$1.25—
claude-opus-4-5$5.00$25.00$0.50$6.25—
claude-opus-4-6$5.00$25.00$0.50$6.25—
claude-opus-4-7$5.00$25.00$0.50$6.25—
claude-opus-4-8$5.00$25.00$0.50$6.25—
claude-opus-5$5.00$25.00$0.50$6.25—
claude-sonnet-4-5$3.00$15.00$0.30$3.75—
claude-sonnet-4-6$3.00$15.00$0.30$3.75—
claude-sonnet-5$3.00$15.00$0.30$3.75—

Google Vertex AI

Google Vertex AI
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gemini-2.5-flash$0.30$2.50$0.03——
gemini-2.5-flash-lite$0.10$0.40$0.01——
gemini-2.5-pro$1.25$10.00$0.125——
Above 200K input tokens$2.50$15.00$0.25——
gemini-3-flash-preview$0.50$3.00$0.05——
gemini-3.1-flash-lite$0.25$1.50$0.025——
gemini-3.1-pro-preview$2.00$12.00$0.20——
Above 200K input tokens$4.00$18.00$0.40——
gemini-3.5-flash$1.50$9.00$0.15——
gemini-3.5-flash-lite$0.30$2.50$0.03——

Azure AI Foundry

Azure AI Foundry
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gpt-5$1.25$10.00$0.125——
gpt-5-mini$0.25$2.00$0.025——
gpt-5-nano$0.05$0.40$0.005——
gpt-5.1$1.25$10.00$0.125——
gpt-5.2$1.75$14.00$0.175——
gpt-5.4$2.50$15.00$0.25——
Above 272K input tokens$5.00$22.50$0.50——
gpt-5.4-mini$0.75$4.50$0.075——
gpt-5.4-nano$0.20$1.25$0.02——
gpt-5.4-pro$30.00$180.00———
Above 272K input tokens$60.00$270.00———
gpt-5.5$5.00$30.00$0.50——
gpt-5.6-luna$1.00$6.00$0.10$1.25—
Above 272K input tokens$2.00$9.00$0.20$2.50—
gpt-5.6-sol$5.00$30.00$0.50$6.25—
Above 272K input tokens$10.00$45.00$1.00$12.50—
gpt-5.6-terra$2.50$15.00$0.25$3.125—
Above 272K input tokens$5.00$22.50$0.50$6.25—

Cloudflare Workers AI

Cloudflare Workers AI
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
deepseek-r1-distill-qwen-32b$0.497$4.881———
gemma-4-26b-a4b-it$0.10$0.30———
gemma-sea-lion-v4-27b-it$0.351$0.555———
glm-4.7-flash$0.06$0.40———
glm-5.2$1.40$4.40$0.26——
gpt-oss-120b$0.35$0.75———
gpt-oss-20b$0.20$0.30———
granite-4.0-h-micro$0.017$0.112———
kimi-k2.5$0.60$3.00$0.10——
kimi-k2.6$0.95$4.00$0.16——
kimi-k2.7-code$0.95$4.00$0.19——
llama-3.1-70b-instruct-fp8-fast$0.293$2.253———
llama-3.1-8b$0.045$0.384———
llama-3.1-8b-instruct-fp8$0.152$0.287———
llama-3.2-11b-vision-instruct$0.049$0.676———
llama-3.2-1b-instruct$0.027$0.201———
llama-3.2-3b-instruct$0.051$0.335———
llama-3.3-70b$0.293$2.253———
llama-4-scout-17b-16e-instruct$0.27$0.85———
llama-guard-3-8b$0.484$0.03———
mistral-7b-instruct-v0.1$0.11$0.19———
mistral-small-3.1$0.351$0.555———
nemotron-3-120b-a12b$0.50$1.50———
qwen2.5-coder-32b-instruct$0.66$1.00———
qwen3-30b$0.051$0.335———
qwq-32b$0.66$1.00———

DeepInfra

DeepInfra
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
deepseek-r1-0528$0.50$2.15$0.35——
deepseek-v3$0.32$0.89———
deepseek-v3.1$0.25$0.95$0.13——
deepseek-v3.2$0.26$0.38$0.13——
deepseek-v4-flash$0.09$0.18$0.018——
deepseek-v4-flash-0731$0.06$0.18$0.015——
deepseek-v4-flash-vision-exp$0.44$1.32$0.014——
deepseek-v4-pro$1.30$2.60$0.10——
gemma-3-12b-it$0.05$0.15———
gemma-3-27b-it$0.08$0.16———
gemma-3-4b-it$0.05$0.10———
gemma-4-26b-a4b-it$0.07$0.34———
gemma-4-31b-it$0.13$0.38———
gemma-4-31b-it-turbo$0.09$0.34$0.05——
gemma-4-31b-it-ultra$0.27$0.76———
kimi-k2.6$0.75$3.50$0.15——
kimi-k3$2.85$14.25$0.285——
llama-3.3-70b-instruct-turbo$0.10$0.32———
llama-4-scout-17b-16e-instruct$0.10$0.30———
llama-guard-4-12b$0.18$0.18———
meta-llama-3.1-70b-instruct-turbo$0.40$0.40———
meta-llama-3.1-8b-instruct-turbo$0.02$0.04———
mistral-nemo-instruct-2407$0.019$0.03———
mistral-small-24b-instruct-2501$0.05$0.08———
mistral-small-3.2-24b-instruct-2506$0.075$0.20———
mythomax-l2-13b$0.40$0.40———
nemotron-3-nano-30b-a3b$0.05$0.20$0.025——
nemotron-content-safety-3.5$0.20$0.20———
nvidia-nemotron-3-super-120b-a12b$0.085$0.40———
nvidia-nemotron-3-ultra-550b-a55b$0.50$2.20$0.10——
nvidia-nemotron-3.5-lightning$0.08$0.20$0.04——
phi-4$0.07$0.14———
qwen2.5-72b-instruct$0.36$0.40———
qwen3-14b$0.12$0.24———
qwen3-235b-a22b-instruct-2507$0.09$0.55———
qwen3-30b-a3b$0.12$0.50———
qwen3-32b$0.08$0.28———
qwen3-coder-480b-a35b-instruct-turbo$0.30$1.00$0.10——
qwen3-max$1.20$6.00$0.24——
qwen3-max-thinking$1.20$6.00$0.24——
qwen3-next-80b-a3b-instruct$0.09$1.10———
qwen3-vl-235b-a22b-instruct$0.20$0.88$0.11——
qwen3-vl-30b-a3b-instruct$0.15$0.60———
qwen3.5-27b$0.26$2.60———
qwen3.5-35b-a3b$0.14$1.00$0.05——
qwen3.5-397b-a17b$0.45$3.00$0.22——
qwen3.5-9b$0.10$0.15———
qwen3.6-27b$0.32$3.20———
qwen3.6-35b-a3b$0.10$0.95———
qwen3.7-max$2.50$7.50$0.50——
qwen3.8-max$1.65$4.951$0.206——