pricing

The number here is the number on your wallet.

Sizing an AI feature usually means multiplying a guess by a tier name and then booking a call. So here is the whole thing on one page. Every model we serve, at the rate it bills, read live from the same catalog our biller reads, refreshed every 60 seconds. Nothing sits behind a form.

We take a markup on tokens. That markup is already inside each rate below, which is why there is no seat price, no monthly minimum and no surcharge per call. One number to plan against and one number to be charged, and they match.

Model rates

Per 1M tokens, input and output priced separately. Live from the model catalog, refreshed every 60s.

ModelInputOutputProvider
fc:ai21/jamba-large-1.7$2.00$8.00ai21
fc:amazon/nova-premier-v1$2.50$12.50amazon
fc:~anthropic/claude-sonnet-latest$2.00$10.00~anthropic
fc:anthropic/claude-fable-5$10.00$50.00anthropic
fc:~anthropic/claude-fable-latest$10.00$50.00~anthropic
fc:anthropic/claude-opus-4$15.00$75.00anthropic
fc:anthropic/claude-opus-4.1$15.00$75.00anthropic
fc:anthropic/claude-opus-4.5$5.00$25.00anthropic
fc:anthropic/claude-opus-4.6$5.00$25.00anthropic
fc:anthropic/claude-opus-4.7$5.00$25.00anthropic
fc:anthropic/claude-opus-4.7-fast$30.00$150.00anthropic
fc:anthropic/claude-opus-4.8$5.00$25.00anthropic
fc:anthropic/claude-opus-4.8-fast$10.00$50.00anthropic
fc:~anthropic/claude-opus-latest$5.00$25.00~anthropic
fc:anthropic/claude-sonnet-4$3.00$15.00anthropic
fc:anthropic/claude-sonnet-4.5$3.00$15.00anthropic
fc:anthropic/claude-sonnet-4.6$3.00$15.00anthropic
fc:anthropic/claude-sonnet-5$2.00$10.00anthropic
fc:anthropic/claude-opus-5$5.00$25.00anthropic
fc:anthropic/claude-opus-5-fast$10.00$50.00anthropic
fc:cohere/command-a$2.50$10.00cohere
fc:cohere/command-r-plus-08-2024$2.50$10.00cohere
fc:~google/gemini-pro-latest$2.00$12.00~google
fc:google/gemini-2.5-pro$1.25$10.00google
fc:google/gemini-2.5-pro-preview-05-06$1.25$10.00google
fc:google/gemini-2.5-pro-preview$1.25$10.00google
fc:google/gemini-3.1-pro-preview$2.00$12.00google
fc:google/gemini-3.1-pro-preview-customtools$2.00$12.00google
fc:google/gemini-3.5-flash$1.50$9.00google
fc:google/gemini-3-pro-image-preview$2.00$12.00google
fc:google/gemini-3-pro-image$2.00$12.00google
fc:~moonshotai/kimi-latest$3.00$15.00~moonshotai
fc:moonshotai/kimi-k3$3.00$15.00moonshotai
fc:~openai/gpt-latest$5.00$30.00~openai
fc:openai/gpt-audio$2.50$10.00openai
fc:openai/gpt-chat-latest$5.00$30.00openai
fc:openai/gpt-4$30.00$60.00openai
fc:openai/gpt-4-turbo$10.00$30.00openai
fc:openai/gpt-4-turbo-preview$10.00$30.00openai
fc:openai/gpt-4.1$2.00$8.00openai
fc:openai/gpt-4o$2.50$10.00openai
fc:openai/gpt-4o-2024-05-13$5.00$15.00openai
fc:openai/gpt-4o-2024-08-06$2.50$10.00openai
fc:openai/gpt-4o-2024-11-20$2.50$10.00openai
fc:openai/gpt-5$1.25$10.00openai
fc:openai/gpt-5-codex$1.25$10.00openai
fc:openai/gpt-5-image$10.00$10.00openai
fc:openai/gpt-5-pro$15.00$120.00openai
fc:openai/gpt-5.1$1.25$10.00openai
fc:openai/gpt-5.1-chat$1.25$10.00openai
fc:openai/gpt-5.1-codex$1.25$10.00openai
fc:openai/gpt-5.1-codex-max$1.25$10.00openai
fc:openai/gpt-5.2$1.75$14.00openai
fc:openai/gpt-5.2-chat$1.75$14.00openai
fc:openai/gpt-5.2-pro$21.00$168.00openai
fc:openai/gpt-5.2-codex$1.75$14.00openai
fc:openai/gpt-5.3-chat$1.75$14.00openai
fc:openai/gpt-5.3-codex$1.75$14.00openai
fc:openai/gpt-5.4$2.50$15.00openai
fc:openai/gpt-5.4-image-2$8.00$15.00openai
fc:openai/gpt-5.4-pro$30.00$180.00openai
fc:openai/gpt-5.5$5.00$30.00openai
fc:openai/gpt-5.5-pro$30.00$180.00openai
fc:openai/gpt-5.6-sol$5.00$30.00openai
fc:openai/gpt-5.6-sol-pro$5.00$30.00openai
fc:openai/o1$15.00$60.00openai
fc:openai/o1-pro$150.00$600.00openai
fc:openai/o3$2.00$8.00openai
fc:openai/o3-deep-research$10.00$40.00openai
fc:openai/o3-pro$20.00$80.00openai
fc:openai/o4-mini-deep-research$2.00$8.00openai
fc:perplexity/sonar-deep-research$2.00$8.00perplexity
fc:perplexity/sonar-pro$3.00$15.00perplexity
fc:perplexity/sonar-pro-search$3.00$15.00perplexity
fc:perplexity/sonar-reasoning-pro$2.00$8.00perplexity
fc:sakana/fugu-ultra$5.00$30.00sakana
fc:aion-labs/aion-2.0$0.80$1.60aion-labs
fc:aion-labs/aion-3.0$3.00$6.00aion-labs
fc:aion-labs/aion-3.0-mini$0.70$1.40aion-labs
fc:aion-labs/aion-rp-llama-3.1-8b$0.80$1.60aion-labs
fc:allenai/olmo-3-32b-think$0.15$0.50allenai
fc:amazon/nova-2-lite-v1$0.30$2.50amazon
fc:amazon/nova-pro-v1$0.80$3.20amazon
fc:~anthropic/claude-haiku-latest$1.00$5.00~anthropic
fc:anthropic/claude-3-haiku$0.25$1.25anthropic
fc:anthropic/claude-haiku-4.5$1.00$5.00anthropic
fc:arcee-ai/trinity-large-thinking$0.22$0.85arcee-ai
fc:arcee-ai/virtuoso-large$0.75$1.20arcee-ai
fc:baidu/ernie-4.5-vl-424b-a47b$0.42$1.25baidu
fc:bytedance-seed/seed-1.6$0.25$2.00bytedance-seed
fc:bytedance-seed/seed-2.0-lite$0.25$2.00bytedance-seed
fc:cohere/command-r-08-2024$0.15$0.60cohere
fc:deepcogito/cogito-v2.1-671b$1.25$1.25deepcogito
fc:deepseek/deepseek-chat$0.20$0.80deepseek
fc:deepseek/deepseek-chat-v3-0324$0.27$1.12deepseek
fc:deepseek/deepseek-chat-v3.1$0.25$0.95deepseek
fc:deepseek/deepseek-v3.1-terminus$0.27$1.00deepseek
fc:deepseek/deepseek-v4-pro$0.43$0.87deepseek
fc:deepseek/deepseek-r1$0.70$2.50deepseek
fc:deepseek/deepseek-r1-0528$0.50$2.15deepseek
fc:deepseek/deepseek-r1-distill-llama-70b$0.80$0.80deepseek
fc:~google/gemini-flash-latest$1.50$7.50~google
fc:google/gemini-2.5-flash$0.30$2.50google
fc:google/gemini-3-flash-preview$0.50$3.00google
fc:google/gemini-3.1-flash-lite$0.25$1.50google
fc:google/gemini-3.1-flash-lite-preview$0.25$1.50google
fc:google/gemini-3.5-flash-lite$0.30$2.50google
fc:google/gemini-3.6-flash$1.50$7.50google
fc:google/gemma-2-27b-it$0.65$0.65google
fc:google/gemini-2.5-flash-image$0.30$2.50google
fc:google/gemini-3.1-flash-image-preview$0.50$3.00google
fc:google/gemini-3.1-flash-image$0.50$3.00google
fc:google/gemini-3.1-flash-lite-image$0.25$1.50google
fc:inception/mercury-2$0.25$0.75inception
fc:inclusionai/ling-2.6-1t$0.07$0.63inclusionai
fc:inclusionai/ring-2.6-1t$0.07$0.63inclusionai
fc:kwaipilot/kat-coder-air-v2.5$0.15$0.60kwaipilot
fc:kwaipilot/kat-coder-pro-v2$0.30$1.20kwaipilot
fc:kwaipilot/kat-coder-pro-v2.5$0.74$2.96kwaipilot
fc:anthracite-org/magnum-v4-72b$3.00$5.00anthracite-org
fc:mancer/weaver$0.50$0.75mancer
fc:meituan/longcat-2.0$0.30$1.20meituan
fc:meta-llama/llama-4-maverick$0.20$0.80meta-llama
fc:meta/muse-spark-1.1$1.25$4.25meta
fc:minimax/minimax-m1$0.55$2.20minimax
fc:minimax/minimax-m2$0.26$1.02minimax
fc:minimax/minimax-m2-her$0.30$1.20minimax
fc:minimax/minimax-m2.1$0.30$1.20minimax
fc:minimax/minimax-m2.5$0.15$0.90minimax
fc:minimax/minimax-m2.7$0.25$1.00minimax
fc:minimax/minimax-m3$0.30$1.20minimax
fc:minimax/minimax-01$0.20$1.10minimax
fc:mistralai/mistral-large$2.00$6.00mistralai
fc:mistralai/mistral-large-2407$2.00$6.00mistralai
fc:mistralai/codestral-2508$0.30$0.90mistralai
fc:mistralai/devstral-2512$0.40$2.00mistralai
fc:mistralai/mistral-large-2512$0.50$1.50mistralai
fc:mistralai/mistral-medium-3$0.40$2.00mistralai
fc:mistralai/mistral-medium-3.1$0.40$2.00mistralai
fc:mistralai/mistral-medium-3-5$1.50$7.50mistralai
fc:mistralai/mistral-small-3.1-24b-instruct$0.35$0.55mistralai
fc:mistralai/mistral-small-2603$0.15$0.60mistralai
fc:mistralai/mixtral-8x22b-instruct$2.00$6.00mistralai
fc:mistralai/mistral-saba$0.20$0.60mistralai
fc:moonshotai/kimi-k2$0.57$2.30moonshotai
fc:moonshotai/kimi-k2-0905$0.60$2.50moonshotai
fc:moonshotai/kimi-k2-thinking$0.60$2.50moonshotai
fc:moonshotai/kimi-k2.5$0.57$2.85moonshotai
fc:moonshotai/kimi-k2.6$0.65$2.72moonshotai
fc:moonshotai/kimi-k2.7-code$0.73$3.50moonshotai
fc:morph/morph-v3-fast$0.80$1.20morph
fc:morph/morph-v3-large$0.90$1.90morph
fc:nex-agi/nex-n2-pro$0.25$1.00nex-agi
fc:nousresearch/hermes-3-llama-3.1-405b$1.00$1.00nousresearch
fc:nousresearch/hermes-3-llama-3.1-70b$0.70$0.70nousresearch
fc:nousresearch/hermes-4-405b$1.00$3.00nousresearch
fc:nvidia/nemotron-3-ultra-550b-a55b$0.50$2.20nvidia
fc:~openai/gpt-mini-latest$0.75$4.50~openai
fc:openai/gpt-audio-mini$0.60$2.40openai
fc:openai/gpt-3.5-turbo$0.50$1.50openai
fc:openai/gpt-3.5-turbo-0613$1.00$2.00openai
fc:openai/gpt-3.5-turbo-16k$3.00$4.00openai
fc:openai/gpt-3.5-turbo-instruct$1.50$2.00openai
fc:openai/gpt-4.1-mini$0.40$1.60openai
fc:openai/gpt-4o-mini$0.15$0.60openai
fc:openai/gpt-4o-mini-2024-07-18$0.15$0.60openai
fc:openai/gpt-5-image-mini$2.50$2.00openai
fc:openai/gpt-5-mini$0.25$2.00openai
fc:openai/gpt-5.1-codex-mini$0.25$2.00openai
fc:openai/gpt-5.4-mini$0.75$4.50openai
fc:openai/gpt-5.4-nano$0.20$1.25openai
fc:openai/gpt-5.6-luna$0.50$3.00openai
fc:openai/gpt-5.6-luna-pro$0.50$3.00openai
fc:openai/gpt-5.6-terra$1.25$7.50openai
fc:openai/gpt-5.6-terra-pro$1.25$7.50openai
fc:openai/o3-mini$1.10$4.40openai
fc:openai/o3-mini-high$1.10$4.40openai
fc:openai/o4-mini$1.10$4.40openai
fc:openai/o4-mini-high$1.10$4.40openai
fc:perceptron/perceptron-mk1$0.15$1.50perceptron
fc:perplexity/sonar$1.00$1.00perplexity
fc:qwen/qwen-plus-2025-07-28$0.26$0.78qwen
fc:qwen/qwen-plus-2025-07-28:thinking$0.26$0.78qwen
fc:qwen/qwen-plus$0.26$0.78qwen
fc:qwen/qwen2.5-vl-72b-instruct$0.80$1.00qwen
fc:qwen/qwen3-14b$0.23$0.91qwen
fc:qwen/qwen3-235b-a22b$0.45$1.82qwen
fc:qwen/qwen3-235b-a22b-2507$0.09$0.55qwen
fc:qwen/qwen3-235b-a22b-thinking-2507$0.30$3.00qwen
fc:qwen/qwen3-30b-a3b$0.12$0.50qwen
fc:qwen/qwen3-30b-a3b-thinking-2507$0.13$1.56qwen
fc:qwen/qwen3-coder$0.30$1.00qwen
fc:qwen/qwen3-coder-flash$0.20$0.97qwen
fc:qwen/qwen3-coder-next$0.11$0.80qwen
fc:qwen/qwen3-coder-plus$0.65$3.25qwen
fc:qwen/qwen3-max$0.78$3.90qwen
fc:qwen/qwen3-max-thinking$0.78$3.90qwen
fc:qwen/qwen3-next-80b-a3b-instruct$0.10$1.10qwen
fc:qwen/qwen3-next-80b-a3b-thinking$0.10$0.78qwen
fc:qwen/qwen3-vl-235b-a22b-instruct$0.21$1.90qwen
fc:qwen/qwen3-vl-235b-a22b-thinking$0.26$2.60qwen
fc:qwen/qwen3-vl-30b-a3b-instruct$0.15$0.60qwen
fc:qwen/qwen3-vl-30b-a3b-thinking$0.13$1.56qwen
fc:qwen/qwen3-vl-8b-thinking$0.12$1.36qwen
fc:qwen/qwen3.5-397b-a17b$0.39$2.34qwen
fc:qwen/qwen3.5-plus-02-15$0.26$1.56qwen
fc:qwen/qwen3.5-plus-20260420$0.30$1.80qwen
fc:qwen/qwen3.5-122b-a10b$0.26$2.08qwen
fc:qwen/qwen3.5-27b$0.20$1.56qwen
fc:qwen/qwen3.5-35b-a3b$0.14$1.00qwen
fc:qwen/qwen3.6-27b$0.30$2.00qwen
fc:qwen/qwen3.6-35b-a3b$0.14$1.00qwen
fc:qwen/qwen3.6-flash$0.19$1.13qwen
fc:qwen/qwen3.6-max-preview$1.04$6.24qwen
fc:qwen/qwen3.6-plus$0.33$1.95qwen
fc:qwen/qwen3.7-max$1.48$4.42qwen
fc:qwen/qwen3.7-plus$0.32$1.28qwen
fc:qwen/qwen-2.5-coder-32b-instruct$0.66$1.00qwen
fc:relace/relace-apply-3$0.85$1.25relace
fc:relace/relace-search$1.00$3.00relace
fc:undi95/remm-slerp-l2-13b$0.45$0.65undi95
fc:sao10k/l3.1-euryale-70b$0.85$0.85sao10k
fc:sao10k/l3.3-euryale-70b$0.65$0.75sao10k
fc:stepfun/step-3.7-flash$0.20$1.15stepfun
fc:tencent/hunyuan-a13b-instruct$0.14$0.57tencent
fc:tencent/hy3$0.13$0.53tencent
fc:thedrummer/cydonia-24b-v4.1$0.30$0.50thedrummer
fc:thedrummer/rocinante-12b$0.25$0.50thedrummer
fc:thedrummer/skyfall-36b-v2$0.55$0.80thedrummer
fc:thinkingmachines/inkling$1.00$4.05thinkingmachines
fc:upstage/solar-pro-3$0.15$0.60upstage
fc:cognitivecomputations/dolphin-mistral-24b-venice-edition$0.20$0.90cognitivecomputations
fc:microsoft/wizardlm-2-8x22b$0.62$0.62microsoft
fc:writer/palmyra-x5$0.60$6.00writer
fc:x-ai/grok-4.20$1.25$2.50x-ai
fc:x-ai/grok-4.20-multi-agent$1.25$2.50x-ai
fc:x-ai/grok-4.3$1.25$2.50x-ai
fc:x-ai/grok-4.5$2.00$6.00x-ai
fc:x-ai/grok-build-0.1$1.00$2.00x-ai
fc:~x-ai/grok-latest$2.00$6.00~x-ai
fc:xiaomi/mimo-v2.5-pro$0.43$0.87xiaomi
fc:z-ai/glm-4.5$0.60$2.20z-ai
fc:z-ai/glm-4.5-air$0.13$0.85z-ai
fc:z-ai/glm-4.5v$0.60$1.80z-ai
fc:z-ai/glm-4.6$0.50$2.00z-ai
fc:z-ai/glm-4.6v$0.30$0.90z-ai
fc:z-ai/glm-4.7$0.40$1.75z-ai
fc:z-ai/glm-5$0.95$2.55z-ai
fc:z-ai/glm-5-turbo$1.20$4.00z-ai
fc:z-ai/glm-5.1$0.97$3.04z-ai
fc:z-ai/glm-5.2$0.75$2.35z-ai
fc:z-ai/glm-5v-turbo$1.20$4.00z-ai
fc:amazon/nova-lite-v1$0.06$0.24amazon
fc:amazon/nova-micro-v1$0.04$0.14amazon
fc:bytedance-seed/seed-1.6-flash$0.07$0.30bytedance-seed
fc:bytedance-seed/seed-2.0-mini$0.10$0.40bytedance-seed
fc:bytedance/ui-tars-1.5-7b$0.10$0.20bytedance
fc:cohere/command-r7b-12-2024$0.04$0.15cohere
fc:deepseek/deepseek-v3.2$0.27$0.40deepseek
fc:deepseek/deepseek-v3.2-exp$0.27$0.41deepseek
fc:deepseek/deepseek-v4-flash$0.14$0.28deepseek
fc:google/gemini-2.5-flash-lite$0.10$0.40google
fc:google/gemma-3-12b-it$0.05$0.15google
fc:google/gemma-3-27b-it$0.08$0.45google
fc:google/gemma-3-4b-it$0.05$0.10google
fc:google/gemma-3n-e4b-it$0.06$0.12google
fc:google/gemma-4-26b-a4b-it$0.14$0.42google
fc:google/gemma-4-31b-it$0.14$0.40google
fc:ibm-granite/granite-4.0-h-micro$0.02$0.11ibm-granite
fc:ibm-granite/granite-4.1-8b$0.05$0.10ibm-granite
fc:inclusionai/ling-2.6-flash$0.01$0.03inclusionai
fc:meta-llama/llama-3.1-70b-instruct$0.40$0.40meta-llama
fc:meta-llama/llama-3.1-8b-instruct$0.05$0.08meta-llama
fc:meta-llama/llama-3.2-1b-instruct$0.03$0.20meta-llama
fc:meta-llama/llama-3.2-3b-instruct$0.05$0.33meta-llama
fc:meta-llama/llama-3.3-70b-instruct$0.13$0.40meta-llama
fc:meta-llama/llama-4-scout$0.10$0.30meta-llama
fc:meta-llama/llama-guard-4-12b$0.18$0.18meta-llama
fc:microsoft/phi-4$0.07$0.14microsoft
fc:mistralai/ministral-14b-2512$0.20$0.20mistralai
fc:mistralai/ministral-3b-2512$0.10$0.10mistralai
fc:mistralai/ministral-8b-2512$0.15$0.15mistralai
fc:mistralai/mistral-nemo$0.02$0.03mistralai
fc:mistralai/mistral-small-24b-instruct-2501$0.05$0.08mistralai
fc:mistralai/mistral-small-3.2-24b-instruct$0.10$0.30mistralai
fc:mistralai/voxtral-small-24b-2507$0.10$0.30mistralai
fc:gryphe/mythomax-l2-13b$0.06$0.06gryphe
fc:nex-agi/nex-n2-mini$0.02$0.10nex-agi
fc:nousresearch/hermes-4-70b$0.13$0.40nousresearch
fc:nvidia/nemotron-3-nano-30b-a3b$0.05$0.20nvidia
fc:nvidia/nemotron-3-super-120b-a12b$0.08$0.40nvidia
fc:openai/gpt-4.1-nano$0.10$0.40openai
fc:openai/gpt-5-nano$0.05$0.40openai
fc:openai/gpt-oss-120b$0.04$0.17openai
fc:openai/gpt-oss-20b$0.03$0.14openai
fc:openai/gpt-oss-safeguard-20b$0.07$0.30openai
fc:poolside/laguna-m.1$0.20$0.40poolside
fc:poolside/laguna-s-2.1$0.10$0.20poolside
fc:poolside/laguna-xs-2.1$0.06$0.12poolside
fc:qwen/qwen-2.5-7b-instruct$0.04$0.10qwen
fc:qwen/qwen3-30b-a3b-instruct-2507$0.05$0.19qwen
fc:qwen/qwen3-32b$0.08$0.28qwen
fc:qwen/qwen3-8b$0.12$0.45qwen
fc:qwen/qwen3-coder-30b-a3b-instruct$0.07$0.27qwen
fc:qwen/qwen3-vl-32b-instruct$0.10$0.42qwen
fc:qwen/qwen3-vl-8b-instruct$0.12$0.45qwen
fc:qwen/qwen3.5-9b$0.10$0.15qwen
fc:qwen/qwen3.5-flash-02-23$0.07$0.26qwen
fc:qwen/qwen-2.5-72b-instruct$0.36$0.40qwen
fc:rekaai/reka-edge$0.10$0.10rekaai
fc:rekaai/reka-flash-3$0.10$0.20rekaai
fc:sao10k/l3-lunaris-8b$0.04$0.05sao10k
fc:stepfun/step-3.5-flash$0.10$0.30stepfun
fc:tencent/hy3-preview$0.06$0.21tencent
fc:thedrummer/unslopnemo-12b$0.40$0.40thedrummer
fc:xiaomi/mimo-v2.5$0.14$0.28xiaomi
fc:z-ai/glm-4.7-flash$0.06$0.40z-ai

317 of 317 models shown · scroll for more

How the price is made

Take a model, take what the provider charges us for it, add our markup, publish the total. That total is the rate in the table and the amount your wallet is debited for the tokens you burned. There is no second line item further down the invoice, because the margin was in the first one.

Two consequences worth knowing before you build a cost model on this. A per-token markup means we earn nothing from you until you make a call, so a quiet month costs you what it should, which is nothing. And it means comparing a Ringside rate against a raw provider list price will show a difference. That difference is the metering, the wallets, the per-tenant limits, the isolation and the margin reporting you would otherwise be running yourself, and you are welcome to price your own engineering time against it.

Topping up

Prepaid wallet through Stripe. Minimum top-up is $10, credits don't expire, and you can hold a balance for as long as you like. Spend draws down as calls complete.

Capping your customers

Set monthly_budget_usd on a customer and calls past it return 402 before any provider spend happens. Per-customer wallets and rpm/tpm limits are on the same object.

Prompt caching

An agent loop re-sends the same system prompt and the same tool list on every single turn. Thirty turns, thirty copies of a few thousand identical tokens, all billed at full price. Caching stops the re-billing. The first call stores that prefix, and every later call starting with the same bytes reads it back at a discount instead of paying to read it fresh.

How big the discount is depends entirely on the provider, and the spread is wide enough to change which model you should be running. We bill the real cached cost and keep only our markup on it, so a warm cache shows up as a smaller number on your wallet rather than a fatter one on ours.

ProviderCached readCache write
Anthropic, DeepSeek10% of input1.25x once
OpenAI, Mistral50% of inputno premium
Google, xAI25% of inputno premium

Read that table before you assume a flat rule. On Anthropic and DeepSeek you pay a one-time 1.25x to warm the cache and then reads land at a tenth of input, which rewards a long-lived prefix. On OpenAI, Mistral, Google and xAI there is no write surcharge at all and the discount simply applies to the cached portion of the read. Providers with no published cache rate bill cached tokens at full input price, so caching costs you nothing there and saves you nothing either.

Caching is opt-in per request. It pays off on long stable prefixes, a fixed system prompt and a tool manifest you re-send every turn. Every response breaks out what was written to cache and what was read, so your hit rate is measured rather than assumed.

Caching in the API docs →

Managed RAG storage

A vector store costs tokens while it is working and storage while it is sitting there. Parse, embed, query and re-embed all bill as tokens at the model rates above. What is left is the footprint, priced per GB-day on the three lines below and metered daily.

Vector index
$0.05
per GB-day

Embeddings dominate the footprint. Chunk text and the inverted index add a small overhead per store.

File storage
$0.04
per GB-day

The raw source files, kept so a re-embed never needs a re-upload and runs can serve content back.

GraphRAG graph
$0.08
per GB-day

Entity and relationship graph pulled out of your corpus for multi-hop retrieval. Sized and free-tiered exactly like the vector index.

First 1 GB-day per store per day is free on both the vector index and the GraphRAG graph. Wallet credits don't expire. A re-embed during a model migration costs embedding tokens only, because the parse step is already cached.

Worked example. A 50 MB PDF corpus, parsed and embedded and indexed, lands around 200 MB on the vector index. That sits inside the daily free allowance, so the vector and graph lines read $0 until a store grows past 1 GB. File storage bills from the first byte. Query tokens are model rates, above.

How managed RAG works →

Before you commit

What about volume discounts?

Applied automatically and reflected on this page. A rate that looks too good is not a mistake, it is the negotiated one.

What if a provider raises prices?

This table follows the catalog within 60 seconds. We would rather show you a rate that moved than hold a stale one you plan against and then get billed past.

Do unused credits expire?

No. Top up $10 today, spend it over six months, nothing evaporates on a renewal date because there is no renewal date.

Are there enterprise plans?

Yes. SLA, dedicated routing and custom retention are all on the table. Sales is the same email as everything else.

Ten dollars is enough to find out

That covers a real integration, a few thousand metered calls and a margin report with your own customers on it. No card on file beyond the top-up, no seat to cancel later.