Pricing

See how much each one of our models cost across different providers.

An indented row under a model is its rate once a request's input passes the size shown.

The two cache write columns are the provider's 5-minute and 1-hour cache rates. A dash means the provider publishes no rate for that line on that route.

OpenAI

OpenAI
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gpt-3.5-turbo$0.50$1.50
gpt-4.1$2.00$8.00$0.50
gpt-4.1-mini$0.40$1.60$0.10
gpt-4.1-nano$0.10$0.40$0.025
gpt-4o$2.50$10.00$1.25
gpt-4o-mini$0.15$0.60$0.075
gpt-5$1.25$10.00$0.125
gpt-5-mini$0.25$2.00$0.025
gpt-5-nano$0.05$0.40$0.005
gpt-5-pro$15.00$120.00
gpt-5.1$1.25$10.00$0.125
gpt-5.2$1.75$14.00$0.175
gpt-5.2-pro$21.00$168.00
gpt-5.4$2.50$15.00$0.25
Above 272K input tokens$5.00$22.50$0.50
gpt-5.4-mini$0.75$4.50$0.075
gpt-5.4-nano$0.20$1.25$0.02
gpt-5.4-pro$30.00$180.00
Above 272K input tokens$60.00$270.00
gpt-5.5$5.00$30.00$0.50
gpt-5.5-pro$30.00$180.00
gpt-5.6$5.00$30.00$0.50$6.25
Above 272K input tokens$10.00$45.00$1.00$12.50
gpt-5.6-luna$1.00$6.00$0.10$1.25
Above 272K input tokens$2.00$9.00$0.20$2.50
gpt-5.6-sol$5.00$30.00$0.50$6.25
Above 272K input tokens$10.00$45.00$1.00$12.50
gpt-5.6-terra$2.50$15.00$0.25$3.125
Above 272K input tokens$5.00$22.50$0.50$6.25
o1$15.00$60.00$7.50
o3$2.00$8.00$0.50
o3-mini$1.10$4.40$0.55
o3-pro$20.00$80.00
o4-mini$1.10$4.40$0.275

Anthropic

Anthropic
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
claude-fable-5$10.00$50.00$1.00$12.50$20.00
claude-haiku-4-5$1.00$5.00$0.10$1.25$2.00
claude-opus-4-5$5.00$25.00$0.50$6.25$10.00
claude-opus-4-6$5.00$25.00$0.50$6.25$10.00
claude-opus-4-7$5.00$25.00$0.50$6.25$10.00
claude-opus-4-8$5.00$25.00$0.50$6.25$10.00
claude-opus-5$5.00$25.00$0.50$6.25$10.00
claude-sonnet-4-5$3.00$15.00$0.30$3.75$6.00
Above 200K input tokens$6.00$22.50$0.60$7.50$12.00
claude-sonnet-4-6$3.00$15.00$0.30$3.75$6.00
claude-sonnet-5$3.00$15.00$0.30$3.75$6.00

Google

Google
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gemini-2.5-flash$0.30$2.50$0.03
gemini-2.5-flash-lite$0.10$0.40$0.01
gemini-2.5-pro$1.25$10.00$0.125
Above 200K input tokens$2.50$15.00$0.25
gemini-3-flash-preview$0.50$3.00$0.05
gemini-3.1-flash-lite$0.25$1.50$0.025
gemini-3.1-pro-preview$2.00$12.00$0.20
Above 200K input tokens$4.00$18.00$0.40
gemini-3.5-flash$1.50$9.00$0.15
gemini-3.6-flash$0.75$3.75$0.075
gemini-3.7-flash$0.75$3.75$0.075

Amazon Bedrock

Amazon Bedrock
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
claude-fable-5$10.00$50.00$1.00$12.50
claude-haiku-4-5$1.00$5.00$0.10$1.25
claude-opus-4-5$5.00$25.00$0.50$6.25
claude-opus-4-6$5.00$25.00$0.50$6.25
claude-opus-4-7$5.00$25.00$0.50$6.25
claude-opus-4-8$5.00$25.00$0.50$6.25
claude-opus-5$5.00$25.00$0.50$6.25
claude-sonnet-4-5$3.00$15.00$0.30$3.75
claude-sonnet-4-6$3.00$15.00$0.30$3.75
claude-sonnet-5$3.00$15.00$0.30$3.75

Google Vertex AI

Google Vertex AI
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gemini-2.5-flash$0.30$2.50$0.03
gemini-2.5-flash-lite$0.10$0.40$0.01
gemini-2.5-pro$1.25$10.00$0.125
Above 200K input tokens$2.50$15.00$0.25
gemini-3-flash-preview$0.50$3.00$0.05
gemini-3.1-flash-lite$0.25$1.50$0.025
gemini-3.1-pro-preview$2.00$12.00$0.20
Above 200K input tokens$4.00$18.00$0.40
gemini-3.5-flash$1.50$9.00$0.15

Azure AI Foundry

Azure AI Foundry
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
gpt-5$1.25$10.00$0.125
gpt-5-mini$0.25$2.00$0.025
gpt-5-nano$0.05$0.40$0.005
gpt-5.1$1.25$10.00$0.125
gpt-5.2$1.75$14.00$0.175
gpt-5.4$2.50$15.00$0.25
Above 272K input tokens$5.00$22.50$0.50
gpt-5.4-mini$0.75$4.50$0.075
gpt-5.4-nano$0.20$1.25$0.02
gpt-5.4-pro$30.00$180.00
Above 272K input tokens$60.00$270.00
gpt-5.5$5.00$30.00$0.50
gpt-5.6-luna$1.00$6.00$0.10$1.25
Above 272K input tokens$2.00$9.00$0.20$2.50
gpt-5.6-sol$5.00$30.00$0.50$6.25
Above 272K input tokens$10.00$45.00$1.00$12.50
gpt-5.6-terra$2.50$15.00$0.25$3.125
Above 272K input tokens$5.00$22.50$0.50$6.25

Cloudflare Workers AI

Cloudflare Workers AI
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
deepseek-r1-distill-qwen-32b$0.497$4.881
gemma-4-26b-a4b-it$0.10$0.30
gemma-sea-lion-v4-27b-it$0.351$0.555
glm-4.7-flash$0.06$0.40
glm-5.2$1.40$4.40$0.26
gpt-oss-120b$0.35$0.75
gpt-oss-20b$0.20$0.30
granite-4.0-h-micro$0.017$0.112
kimi-k2.5$0.60$3.00$0.10
kimi-k2.6$0.95$4.00$0.16
kimi-k2.7-code$0.95$4.00$0.19
llama-3.1-70b-instruct-fp8-fast$0.293$2.253
llama-3.1-8b$0.045$0.384
llama-3.1-8b-instruct-fp8$0.152$0.287
llama-3.2-11b-vision-instruct$0.049$0.676
llama-3.2-1b-instruct$0.027$0.201
llama-3.2-3b-instruct$0.051$0.335
llama-3.3-70b$0.293$2.253
llama-4-scout-17b-16e-instruct$0.27$0.85
llama-guard-3-8b$0.484$0.03
mistral-7b-instruct-v0.1$0.11$0.19
mistral-small-3.1$0.351$0.555
nemotron-3-120b-a12b$0.50$1.50
qwen2.5-coder-32b-instruct$0.66$1.00
qwen3-30b$0.051$0.335
qwq-32b$0.66$1.00

DeepInfra

DeepInfra
ModelInput $ / M TokensOutput $ / M TokensCache read $ / M TokensCache write (5 min) $ / M TokensCache write (1 hour) $ / M Tokens
deepseek-r1-0528$0.50$2.15$0.35
deepseek-v3$0.32$0.89
deepseek-v3-0324$0.24$0.90$0.135
deepseek-v3.1$0.25$0.95$0.13
deepseek-v3.2$0.26$0.38$0.13
deepseek-v4-flash$0.09$0.18$0.018
deepseek-v4-flash-0731$0.08$0.18$0.016
deepseek-v4-pro$1.30$2.60$0.10
gemma-3-12b-it$0.05$0.15
gemma-3-27b-it$0.08$0.16
gemma-3-4b-it$0.05$0.10
gemma-4-26b-a4b-it$0.07$0.34
gemma-4-31b-it$0.13$0.38
gemma-4-31b-it-turbo$0.09$0.34$0.05
gemma-4-31b-it-ultra$0.27$0.76
gemma-4-e4b-it$0.02$0.10
kimi-k2.5$0.45$2.25$0.07
kimi-k2.6$0.75$3.50$0.15
kimi-k2.7-code$0.68$3.40$0.136
kimi-k3$2.85$14.25$0.285
llama-3.3-70b-instruct-turbo$0.10$0.32
llama-4-maverick-17b-128e-instruct-fp8$0.20$0.80
llama-4-scout-17b-16e-instruct$0.10$0.30
llama-guard-4-12b$0.18$0.18
meta-llama-3.1-70b-instruct-turbo$0.40$0.40
meta-llama-3.1-8b-instruct-turbo$0.02$0.04
mistral-nemo-instruct-2407$0.019$0.03
mistral-small-24b-instruct-2501$0.05$0.08
mistral-small-3.2-24b-instruct-2506$0.075$0.20
mythomax-l2-13b$0.40$0.40
nemotron-3-nano-30b-a3b$0.05$0.20$0.025
nemotron-content-safety-3.5$0.20$0.20
nvidia-nemotron-3-super-120b-a12b$0.085$0.40
nvidia-nemotron-3-ultra-550b-a55b$0.50$2.20$0.10
nvidia-nemotron-3.5-lightning$0.08$0.20$0.04
phi-4$0.07$0.14
qwen2.5-72b-instruct$0.36$0.40
qwen3-14b$0.12$0.24
qwen3-235b-a22b-instruct-2507$0.09$0.55
qwen3-30b-a3b$0.12$0.50
qwen3-32b$0.08$0.28
qwen3-coder-480b-a35b-instruct-turbo$0.30$1.00$0.10
qwen3-max$1.20$6.00$0.24
qwen3-max-thinking$1.20$6.00$0.24
qwen3-next-80b-a3b-instruct$0.09$1.10
qwen3-vl-235b-a22b-instruct$0.20$0.88$0.11
qwen3-vl-30b-a3b-instruct$0.15$0.60
qwen3.5-122b-a10b$0.29$2.40
qwen3.5-27b$0.26$2.60
qwen3.5-35b-a3b$0.14$1.00$0.05
qwen3.5-397b-a17b$0.45$3.00$0.22
qwen3.5-9b$0.10$0.15
qwen3.6-27b$0.32$3.20
qwen3.6-35b-a3b$0.10$0.95
qwen3.7-max$2.50$7.50$0.50
qwen3.8-max$1.65$4.951$0.206