live · updated hourly

The AI Price Index

List prices for every language model with published per-token pricing, across every major provider, in one table — plus the two columns nobody else publishes: how much more an output token costs than an input one, and what caching actually saves you.

What the catalogue says today

Every figure below is computed from the live catalogue at load, not typed in by hand.

316
models with published per-token pricing.
across 51 providers
4.0×
median output premium — an output token costs this much more than an input one.
prefill vs decode, priced
90%
median discount on a cached input token, where caching is offered.
177 models publish a cache rate
$0.015
cheapest blended cost per 1M tokens at a 3:1 input:output mix.
inclusionai/ling-2.6-flash

The output premium is the single most useful number on this page. Prefill is compute-bound and parallel; decode is memory-bandwidth-bound and serial. That physical asymmetry is why output costs multiples of input — and why an architecture that generates fewer, shorter outputs beats one that merely switches to a cheaper model.

Every priced model, cheapest blended first

Sorted by blended cost at a 3:1 input:output ratio — the common benchmarking convention. The 10:1 column is closer to what agentic and RAG-heavy traffic actually looks like, and the two orders differ more than most people expect.

Updated Aug 4, 2026, 10:57 PM UTC · source: the public OpenRouter model catalogue.

ModelIn / 1MOut / 1MOutput premiumCache savesBlended 3:1Blended 10:1Context
inclusionAI: Ling-2.6-flash
inclusionai/ling-2.6-flash
$0.010$0.0303.0×80%$0.015$0.012262K
Mistral: Mistral Nemo
mistralai/mistral-nemo
$0.019$0.0301.6×$0.022$0.020131K
IBM: Granite 4.0 Micro
ibm-granite/granite-4.0-h-micro
$0.017$0.1126.6×$0.041$0.026131K
Sao10K: Llama 3 8B Lunaris
sao10k/l3-lunaris-8b
$0.040$0.0501.2×$0.042$0.0418K
Nex AGI: Nex-N2-Mini
nex-agi/nex-n2-mini
$0.025$0.1004.0×90%$0.044$0.032262K
Qwen: Qwen3.7 Flash
qwen/qwen3.7-flash
$0.030$0.1304.3×80%$0.055$0.0391M
OpenAI: gpt-oss-20b
openai/gpt-oss-20b
$0.030$0.1304.3×$0.055$0.039131K
Mistral: Mistral Small 3
mistralai/mistral-small-24b-instruct-2501
$0.050$0.0801.6×$0.057$0.05333K
Meta: Llama 3.1 8B Instruct
meta-llama/llama-3.1-8b-instruct
$0.050$0.0801.6×50%$0.057$0.053131K
Amazon: Nova Micro 1.0
amazon/nova-micro-v1
$0.035$0.1404.0×$0.061$0.045128K
IBM: Granite 4.1 8B
ibm-granite/granite-4.1-8b
$0.050$0.1002.0×$0.063$0.055131K
Google: Gemma 3 4B
google/gemma-3-4b-it
$0.050$0.1002.0×$0.063$0.055131K
Cohere: Command R7B (12-2024)
cohere/command-r7b-12-2024
$0.037$0.1504.0×$0.066$0.048128K
OpenAI: gpt-oss-120b
openai/gpt-oss-120b
$0.037$0.1704.6×$0.070$0.049131K
Meta: Llama 3.2 1B Instruct
meta-llama/llama-3.2-1b-instruct
$0.027$0.2017.4×$0.071$0.04360K
Poolside: Laguna XS 2.1
poolside/laguna-xs-2.1
$0.060$0.1202.0×50%$0.075$0.065262K
Google: Gemma 3n 4B
google/gemma-3n-e4b-it
$0.060$0.1202.0×$0.075$0.06533K
Google: Gemma 3 12B
google/gemma-3-12b-it
$0.050$0.1503.0×$0.075$0.059131K
Qwen: Qwen3 30B A3B Instruct 2507
qwen/qwen3-30b-a3b-instruct-2507
$0.048$0.1934.0×$0.084$0.061262K
NVIDIA: Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
$0.050$0.2004.0×40%$0.087$0.064262K
MythoMax 13B
gryphe/mythomax-l2-13b
$0.080$0.1101.4×$0.087$0.0838K
Microsoft: Phi 4
microsoft/phi-4
$0.070$0.1402.0×$0.088$0.07616K
Tencent: Hy3 preview
tencent/hy3-preview
$0.063$0.2103.3×67%$0.100$0.076262K
Reka Edge
rekaai/reka-edge
$0.100$0.1001.0×$0.100$0.10016K
Mistral: Ministral 3 3B 2512
mistralai/ministral-3b-2512
$0.100$0.1001.0×90%$0.100$0.100131K
Amazon: Nova Lite 1.0
amazon/nova-lite-v1
$0.060$0.2404.0×$0.105$0.076300K
Qwen: Qwen3.5-9B
qwen/qwen3.5-9b
$0.100$0.1501.5×$0.112$0.105262K
DeepSeek V4 Flash Latest
~deepseek/deepseek-v4-flash-latest
$0.090$0.1802.0×80%$0.113$0.0981.0M
DeepSeek: DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
$0.090$0.1802.0×80%$0.113$0.0981.0M
Poolside: Laguna S 2.1
poolside/laguna-s-2.1
$0.090$0.1802.0×90%$0.113$0.0981.0M
Qwen: Qwen3.5-Flash
qwen/qwen3.5-flash-02-23
$0.065$0.2604.0×$0.114$0.0831M
Meta: Llama 3.2 3B Instruct
meta-llama/llama-3.2-3b-instruct
$0.050$0.3306.6×$0.120$0.075131K
Qwen: Qwen3 Coder 30B A3B Instruct
qwen/qwen3-coder-30b-a3b-instruct
$0.070$0.2703.9×$0.120$0.088262K
ByteDance: UI-TARS 7B
bytedance/ui-tars-1.5-7b
$0.100$0.2002.0×$0.125$0.109128K
Reka Flash 3
rekaai/reka-flash-3
$0.100$0.2002.0×$0.125$0.10966K
Qwen: Qwen2.5 7B Instruct
qwen/qwen-2.5-7b-instruct
$0.100$0.2002.0×$0.125$0.10933K
Qwen: Qwen3 32B
qwen/qwen3-32b
$0.080$0.2803.5×$0.130$0.098131K
ByteDance Seed: Seed 1.6 Flash
bytedance-seed/seed-1.6-flash
$0.075$0.3004.0×$0.131$0.095262K
OpenAI: gpt-oss-safeguard-20b
openai/gpt-oss-safeguard-20b
$0.075$0.3004.0×50%$0.131$0.095131K
Mistral: Mistral Small 3.2 24B
mistralai/mistral-small-3.2-24b-instruct
$0.094$0.2502.7×$0.133$0.108256K
OpenAI: GPT-5 Nano
openai/gpt-5-nano
$0.050$0.4008.0×90%$0.137$0.082400K
Google: Gemma 4 26B A4B
google/gemma-4-26b-a4b-it
$0.070$0.3404.9×$0.138$0.095262K
Z.ai: GLM 4.7 Flash
z-ai/glm-4.7-flash
$0.060$0.4006.7×83%$0.145$0.091203K
StepFun: Step 3.5 Flash
stepfun/step-3.5-flash
$0.100$0.3003.0×$0.150$0.118262K
Mistral: Ministral 3 8B 2512
mistralai/ministral-8b-2512
$0.150$0.1501.0×90%$0.150$0.150262K
Mistral: Voxtral Small 24B 2507
mistralai/voxtral-small-24b-2507
$0.100$0.3003.0×90%$0.150$0.11832K
Meta: Llama 4 Scout
meta-llama/llama-4-scout
$0.100$0.3003.0×$0.150$0.1181.3M
Meta: Llama 3.3 70B Instruct
meta-llama/llama-3.3-70b-instruct
$0.100$0.3203.2×$0.155$0.120131K
Google: Gemma 4 31B
google/gemma-4-31b-it
$0.100$0.3403.4×$0.160$0.122262K
NVIDIA: Nemotron 3 Super
nvidia/nemotron-3-super-120b-a12b
$0.085$0.4004.7×$0.164$0.1141M
Google: Gemma 3 27B
google/gemma-3-27b-it
$0.080$0.4505.6×50%$0.172$0.114262K
ByteDance Seed: Seed-2.0-Mini
bytedance-seed/seed-2.0-mini
$0.100$0.4004.0×$0.175$0.127262K
Google: Gemini 2.5 Flash Lite
google/gemini-2.5-flash-lite
$0.100$0.4004.0×90%$0.175$0.1271.0M
OpenAI: GPT-4.1 Nano
openai/gpt-4.1-nano
$0.100$0.4004.0×75%$0.175$0.1271.0M
DeepSeek: DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash
$0.140$0.2802.0×80%$0.175$0.1531.0M
Xiaomi: MiMo-V2.5
xiaomi/mimo-v2.5
$0.140$0.2802.0×98%$0.175$0.1531.1M
Meta: Llama Guard 4 12B
meta-llama/llama-guard-4-12b
$0.180$0.1801.0×$0.180$0.1801.0M
Qwen: Qwen3 VL 32B Instruct
qwen/qwen3-vl-32b-instruct
$0.104$0.4164.0×$0.182$0.132131K
Nous: Hermes 4 70B
nousresearch/hermes-4-70b
$0.130$0.4003.1×$0.198$0.155131K
Mistral: Ministral 3 14B 2512
mistralai/ministral-14b-2512
$0.200$0.2001.0×90%$0.200$0.200262K
Qwen: Qwen3 VL 8B Instruct
qwen/qwen3-vl-8b-instruct
$0.117$0.4553.9×$0.202$0.148262K
Qwen: Qwen3 8B
qwen/qwen3-8b
$0.117$0.4553.9×$0.202$0.148131K
inclusionAI: Ring-2.6-1T
inclusionai/ring-2.6-1t
$0.075$0.6258.3×80%$0.212$0.125262K
inclusionAI: Ling-2.6-1T
inclusionai/ling-2.6-1t
$0.075$0.6258.3×80%$0.212$0.125262K
Qwen: Qwen3 30B A3B
qwen/qwen3-30b-a3b
$0.120$0.5004.2×$0.215$0.155131K
OpenAI: GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro
$0.100$0.6006.0×90%$0.225$0.1451.1M
OpenAI: GPT-5.6 Luna
openai/gpt-5.6-luna
$0.100$0.6006.0×90%$0.225$0.1451.1M
Tencent: Hy3
tencent/hy3
$0.132$0.5284.0×75%$0.231$0.168262K
AllenAI: Olmo 3 32B Think
allenai/olmo-3-32b-think
$0.150$0.5003.3×$0.237$0.18266K
Tencent: Hunyuan A13B Instruct
tencent/hunyuan-a13b-instruct
$0.140$0.5704.1×$0.248$0.179131K
Qwen: Qwen3 235B A22B Instruct 2507
qwen/qwen3-235b-a22b-2507
$0.150$0.5984.0×$0.262$0.190262K
Kwaipilot: KAT-Coder-Air V2.5
kwaipilot/kat-coder-air-v2.5
$0.150$0.6004.0×80%$0.262$0.191256K
Mistral: Mistral Small 4
mistralai/mistral-small-2603
$0.150$0.6004.0×90%$0.262$0.191262K
Upstage: Solar Pro 3
upstage/solar-pro-3
$0.150$0.6004.0×90%$0.262$0.191128K
Qwen: Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b-instruct
$0.150$0.6004.0×$0.262$0.191262K
Cohere: Command R (08-2024)
cohere/command-r-08-2024
$0.150$0.6004.0×$0.262$0.191128K
OpenAI: GPT-4o-mini
openai/gpt-4o-mini
$0.150$0.6004.0×50%$0.262$0.191128K
OpenAI: GPT-4o-mini (2024-07-18)
openai/gpt-4o-mini-2024-07-18
$0.150$0.6004.0×50%$0.262$0.191128K
Qwen: Qwen3 Coder Next
qwen/qwen3-coder-next
$0.120$0.8006.7×42%$0.290$0.182262K
Mistral: Saba
mistralai/mistral-saba
$0.200$0.6003.0×90%$0.300$0.23633K
DeepSeek: DeepSeek V3.2
deepseek/deepseek-v3.2
$0.269$0.4001.5×50%$0.302$0.281164K
DeepSeek: DeepSeek V3.2 Exp
deepseek/deepseek-v3.2-exp
$0.270$0.4101.5×$0.305$0.283164K
Z.ai: GLM 4.5 Air
z-ai/glm-4.5-air
$0.130$0.8506.5×81%$0.310$0.195131K
TheDrummer: Rocinante 12B
thedrummer/rocinante-12b
$0.250$0.5002.0×$0.313$0.27366K
MiniMax: MiniMax M2.5
minimax/minimax-m2.5
$0.150$0.9006.0×67%$0.337$0.218205K
Qwen: Qwen3 Next 80B A3B Instruct
qwen/qwen3-next-80b-a3b-instruct
$0.090$1.1012.2×$0.343$0.182262K
TheDrummer: Cydonia 24B V4.1
thedrummer/cydonia-24b-v4.1
$0.300$0.5001.7×50%$0.350$0.318131K
Meta: Llama 4 Maverick
meta-llama/llama-4-maverick
$0.200$0.8004.0×$0.350$0.2551.0M
Qwen: Qwen3.6 35B A3B
qwen/qwen3.6-35b-a3b
$0.140$1.007.1×$0.355$0.218262K
Qwen: Qwen3.5-35B-A3B
qwen/qwen3.5-35b-a3b
$0.140$1.007.1×$0.355$0.218262K
Qwen2.5 72B Instruct
qwen/qwen-2.5-72b-instruct
$0.360$0.4001.1×$0.370$0.36433K
Inception: Mercury 2
inception/mercury-2
$0.250$0.7503.0×90%$0.375$0.295128K
Venice: Uncensored
cognitivecomputations/dolphin-mistral-24b-venice-edition
$0.200$0.9004.5×$0.375$0.264128K
Qwen: Qwen2.5 VL 72B Instruct
qwen/qwen2.5-vl-72b-instruct
$0.250$0.7503.0×$0.375$0.295128K
Arcee AI: Trinity Large Thinking
arcee-ai/trinity-large-thinking
$0.220$0.8503.9×73%$0.378$0.277262K
Qwen: Qwen3 Coder Flash
qwen/qwen3-coder-flash
$0.195$0.9755.0×80%$0.390$0.2661M
Qwen: Qwen Plus 0728
qwen/qwen-plus-2025-07-28
$0.260$0.7803.0×$0.390$0.3071M
Qwen: Qwen-Plus
qwen/qwen-plus
$0.260$0.7803.0×80%$0.390$0.3071M
Qwen: Qwen3 14B
qwen/qwen3-14b
$0.227$0.9104.0×$0.398$0.290131K
TheDrummer: UnslopNemo 12B
thedrummer/unslopnemo-12b
$0.400$0.4001.0×$0.400$0.4001.0M
Meta: Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct
$0.400$0.4001.0×$0.400$0.400131K
Mistral: Mistral Small 3.1 24B
mistralai/mistral-small-3.1-24b-instruct
$0.351$0.5551.6×$0.402$0.370128K
Qwen: Qwen3 Next 80B A3B Thinking
qwen/qwen3-next-80b-a3b-thinking
$0.150$1.208.0×$0.412$0.245262K
Qwen: Qwen3.6 Flash
qwen/qwen3.6-flash
$0.188$1.136.0×$0.422$0.2731M
DeepSeek: DeepSeek V3.1
deepseek/deepseek-chat-v3.1
$0.250$0.9503.8×48%$0.425$0.314164K
MiniMax: MiniMax-01
minimax/minimax-01
$0.200$1.105.5×$0.425$0.2821.0M
Nex AGI: Nex-N2-Pro
nex-agi/nex-n2-pro
$0.250$1.004.0×90%$0.438$0.318262K
StepFun: Step 3.7 Flash
stepfun/step-3.7-flash
$0.200$1.155.8×80%$0.438$0.286262K
MiniMax: MiniMax M2
minimax/minimax-m2
$0.255$1.024.0×$0.446$0.325205K
Z.ai: GLM 4.6V
z-ai/glm-4.6v
$0.300$0.9003.0×82%$0.450$0.355131K
Mistral: Codestral 2508
mistralai/codestral-2508
$0.300$0.9003.0×90%$0.450$0.355256K
DeepSeek: DeepSeek V3
deepseek/deepseek-chat
$0.257$1.034.0×$0.450$0.328164K
DeepSeek: DeepSeek V3.1 Terminus
deepseek/deepseek-v3.1-terminus
$0.270$1.003.7×50%$0.453$0.336164K
OpenAI: GPT-5.4 Nano
openai/gpt-5.4-nano
$0.200$1.256.3×90%$0.463$0.295400K
MiniMax: MiniMax M2.7
minimax/minimax-m2.7
$0.270$1.084.0×80%$0.473$0.344205K
Qwen: Qwen3 Coder 480B A35B
qwen/qwen3-coder
$0.300$1.003.3×67%$0.475$0.364262K
DeepSeek: DeepSeek V3 0324
deepseek/deepseek-chat-v3-0324
$0.270$1.124.1×50%$0.483$0.347164K
Perceptron: Perceptron Mk1
perceptron/perceptron-mk1
$0.150$1.5010.0×$0.487$0.27333K
Anthropic: Claude 3 Haiku
anthropic/claude-3-haiku
$0.250$1.255.0×88%$0.500$0.341200K
ReMM SLERP 13B
undi95/remm-slerp-l2-13b
$0.450$0.6501.4×$0.500$0.4686K
Meituan: LongCat 2.0
meituan/longcat-2.0
$0.300$1.204.0×98%$0.525$0.3821.0M
MiniMax: MiniMax M3
minimax/minimax-m3
$0.300$1.204.0×80%$0.525$0.3821.0M
Kwaipilot: KAT-Coder-Pro V2
kwaipilot/kat-coder-pro-v2
$0.300$1.204.0×80%$0.525$0.382262K
MiniMax: MiniMax M2-her
minimax/minimax-m2-her
$0.300$1.204.0×90%$0.525$0.38266K
MiniMax: MiniMax M2.1
minimax/minimax-m2.1
$0.300$1.204.0×90%$0.525$0.382205K
Qwen: Qwen3.5-27B
qwen/qwen3.5-27b
$0.195$1.568.0×$0.536$0.319262K
DeepSeek: DeepSeek V4 Pro
deepseek/deepseek-v4-pro
$0.435$0.8702.0×99%$0.544$0.4751.0M
Xiaomi: MiMo-V2.5-Pro
xiaomi/mimo-v2.5-pro
$0.435$0.8702.0×99%$0.544$0.4751.1M
Qwen: Qwen3.7 Plus
qwen/qwen3.7-plus
$0.320$1.284.0×80%$0.560$0.4071M
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
google/gemini-3.1-flash-lite-image
$0.250$1.506.0×$0.563$0.36466K
Google: Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite
$0.250$1.506.0×90%$0.563$0.3641.0M
Google: Gemini 3.1 Flash Lite Preview
google/gemini-3.1-flash-lite-preview
$0.250$1.506.0×90%$0.563$0.3641.0M
Mancer: Weaver (alpha)
mancer/weaver
$0.500$0.7501.5×$0.563$0.5238K
Qwen: Qwen3.5 Plus 2026-02-15
qwen/qwen3.5-plus-02-15
$0.260$1.566.0×$0.585$0.3781M
Qwen: Qwen Plus 0728 (thinking)
qwen/qwen-plus-2025-07-28:thinking
$0.400$1.203.0×$0.600$0.4731M
TheDrummer: Skyfall 36B V2
thedrummer/skyfall-36b-v2
$0.550$0.8001.5×55%$0.613$0.57333K
WizardLM-2 8x22B
microsoft/wizardlm-2-8x22b
$0.620$0.6201.0×$0.620$0.62066K
Baidu: ERNIE 4.5 VL 424B A47B
baidu/ernie-4.5-vl-424b-a47b
$0.420$1.253.0×$0.627$0.495123K
Qwen: Qwen3 VL 235B A22B Instruct
qwen/qwen3-vl-235b-a22b-instruct
$0.210$1.909.0×52%$0.632$0.364262K
Google: Gemma 2 27B
google/gemma-2-27b-it
$0.650$0.6501.0×$0.650$0.6508K
Qwen: Qwen3 VL 8B Thinking
qwen/qwen3-vl-8b-thinking
$0.180$2.1011.7×$0.660$0.355131K
Qwen: Qwen3.5 Plus 2026-04-20
qwen/qwen3.5-plus-20260420
$0.300$1.806.0×$0.675$0.4361M
Thinking Machines: Inkling Small
thinkingmachines/inkling-small
$0.500$1.202.4×80%$0.675$0.564524K
Sao10K: Llama 3.3 Euryale 70B
sao10k/l3.3-euryale-70b
$0.650$0.7501.2×$0.675$0.659131K
ByteDance Seed: Seed-2.0-Lite
bytedance-seed/seed-2.0-lite
$0.250$2.008.0×$0.688$0.409262K
ByteDance Seed: Seed 1.6
bytedance-seed/seed-1.6
$0.250$2.008.0×$0.688$0.409262K
OpenAI: GPT-5.1-Codex-Mini
openai/gpt-5.1-codex-mini
$0.250$2.008.0×88%$0.688$0.409400K
OpenAI: GPT-5 Mini
openai/gpt-5-mini
$0.250$2.008.0×90%$0.688$0.409400K
OpenAI: GPT-4.1 Mini
openai/gpt-4.1-mini
$0.400$1.604.0×75%$0.700$0.5091.0M
Nous: Hermes 3 70B Instruct
nousresearch/hermes-3-llama-3.1-70b
$0.700$0.7001.0×$0.700$0.700131K
Qwen: Qwen3.5-122B-A10B
qwen/qwen3.5-122b-a10b
$0.260$2.088.0×$0.715$0.425262K
Qwen: Qwen3.6 Plus
qwen/qwen3.6-plus
$0.325$1.956.0×$0.731$0.4731M
Z.ai: GLM 4.7
z-ai/glm-4.7
$0.400$1.754.4×80%$0.738$0.523205K
Qwen2.5 Coder 32B Instruct
qwen/qwen-2.5-coder-32b-instruct
$0.660$1.001.5×$0.745$0.69133K
Qwen: Qwen3 235B A22B Thinking 2507
qwen/qwen3-235b-a22b-thinking-2507
$0.230$2.3010.0×$0.747$0.418262K
Mistral: Mistral Large 3 2512
mistralai/mistral-large-2512
$0.500$1.503.0×90%$0.750$0.591262K
Qwen: Qwen3 VL 30B A3B Thinking
qwen/qwen3-vl-30b-a3b-thinking
$0.200$2.4012.0×$0.750$0.400262K
Qwen: Qwen3 30B A3B Thinking 2507
qwen/qwen3-30b-a3b-thinking-2507
$0.200$2.4012.0×$0.750$0.40082K
OpenAI: GPT-3.5 Turbo
openai/gpt-3.5-turbo
$0.500$1.503.0×$0.750$0.59116K
Qwen: Qwen3 235B A22B
qwen/qwen3-235b-a22b
$0.455$1.824.0×$0.796$0.579131K
DeepSeek: R1 Distill Llama 70B
deepseek/deepseek-r1-distill-llama-70b
$0.800$0.8001.0×$0.800$0.8008K
Mistral: Mistral Medium 3.1
mistralai/mistral-medium-3.1
$0.400$2.005.0×90%$0.800$0.545131K
Mistral: Mistral Medium 3
mistralai/mistral-medium-3
$0.400$2.005.0×90%$0.800$0.545131K
Qwen: Qwen3.6 27B
qwen/qwen3.6-27b
$0.289$2.408.3×$0.817$0.481262K
Google: Gemini 3.5 Flash Lite
google/gemini-3.5-flash-lite
$0.300$2.508.3×90%$0.850$0.5001.0M
Amazon: Nova 2 Lite
amazon/nova-2-lite-v1
$0.300$2.508.3×$0.850$0.5001M
Google: Nano Banana (Gemini 2.5 Flash Image)
google/gemini-2.5-flash-image
$0.300$2.508.3×90%$0.850$0.50033K
Google: Gemini 2.5 Flash
google/gemini-2.5-flash
$0.300$2.508.3×90%$0.850$0.5001.0M
Sao10K: Llama 3.1 Euryale 70B v2.2
sao10k/l3.1-euryale-70b
$0.850$0.8501.0×$0.850$0.850131K
Arcee AI: Virtuoso Large
arcee-ai/virtuoso-large
$0.750$1.201.6×$0.863$0.791131K
AionLabs: Aion-3.0-Mini
aion-labs/aion-3.0-mini
$0.700$1.402.0×74%$0.875$0.764131K
Z.ai: GLM 4.6
z-ai/glm-4.6
$0.500$2.004.0×80%$0.875$0.636205K
Qwen: Qwen3.5 397B A17B
qwen/qwen3.5-397b-a17b
$0.390$2.346.0×$0.877$0.567262K
Z.ai: GLM 4.5V
z-ai/glm-4.5v
$0.600$1.803.0×82%$0.900$0.70966K
Morph: Morph V3 Fast
morph/morph-v3-fast
$0.800$1.201.5×$0.900$0.83682K
DeepSeek: R1 0528
deepseek/deepseek-r1-0528
$0.500$2.154.3×30%$0.913$0.650164K
Relace: Relace Apply 3
relace/relace-apply-3
$0.850$1.251.5×$0.950$0.886256K
MiniMax: MiniMax M1
minimax/minimax-m1
$0.550$2.204.0×$0.963$0.7001M
AionLabs: Aion-2.0
aion-labs/aion-2.0
$0.800$1.602.0×75%$1.00$0.873131K
Z.ai: GLM 4.5
z-ai/glm-4.5
$0.600$2.203.7×82%$1.00$0.745131K
AionLabs: Aion-RP 1.0 (8B)
aion-labs/aion-rp-llama-3.1-8b
$0.800$1.602.0×$1.00$0.87333K
Perplexity: Sonar
perplexity/sonar
$1.00$1.001.0×$1.00$1.00127K
Nous: Hermes 3 405B Instruct
nousresearch/hermes-3-llama-3.1-405b
$1.00$1.001.0×$1.00$1.00131K
MoonshotAI: Kimi K2 0711
moonshotai/kimi-k2
$0.570$2.304.0×$1.00$0.727131K
OpenAI: GPT Audio Mini
openai/gpt-audio-mini
$0.600$2.404.0×$1.05$0.764128K
MoonshotAI: Kimi K2.6
moonshotai/kimi-k2.6
$0.589$2.484.2×83%$1.06$0.761262K
MoonshotAI: Kimi K2 Thinking
moonshotai/kimi-k2-thinking
$0.600$2.504.2×75%$1.07$0.773262K
MoonshotAI: Kimi K2 0905
moonshotai/kimi-k2-0905
$0.600$2.504.2×$1.07$0.773262K
Google: Nano Banana 2 (Gemini 3.1 Flash Image)
google/gemini-3.1-flash-image
$0.500$3.006.0×$1.13$0.727131K
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
google/gemini-3.1-flash-image-preview
$0.500$3.006.0×$1.13$0.72766K
Google: Gemini 3 Flash Preview
google/gemini-3-flash-preview
$0.500$3.006.0×90%$1.13$0.7271.0M
MoonshotAI: Kimi K2.5
moonshotai/kimi-k2.5
$0.570$2.855.0×83%$1.14$0.777262K
Morph: Morph V3 Large
morph/morph-v3-large
$0.900$1.902.1×$1.15$0.991262K
DeepSeek: R1
deepseek/deepseek-r1
$0.700$2.503.6×$1.15$0.864164K
Z.ai: GLM 5.2
z-ai/glm-5.2
$0.760$2.423.2×82%$1.18$0.9111.0M
SpaceXAI: Grok Build 0.1
x-ai/grok-build-0.1
$1.00$2.002.0×80%$1.25$1.09256K
Deep Cogito: Cogito v2.1 671B
deepcogito/cogito-v2.1-671b
$1.25$1.251.0×$1.25$1.25128K
OpenAI: GPT-3.5 Turbo (older v0613)
openai/gpt-3.5-turbo-0613
$1.00$2.002.0×$1.25$1.094K
Kwaipilot: KAT-Coder-Pro V2.5
kwaipilot/kat-coder-pro-v2.5
$0.740$2.964.0×80%$1.29$0.942256K
Qwen: Qwen3 Coder Plus
qwen/qwen3-coder-plus
$0.650$3.255.0×80%$1.30$0.8861M
NVIDIA: Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
$0.600$3.606.0×67%$1.35$0.873512K
Z.ai: GLM 5
z-ai/glm-5
$0.950$2.552.7×79%$1.35$1.10205K
Amazon: Nova Pro 1.0
amazon/nova-pro-v1
$0.800$3.204.0×$1.40$1.02300K
MoonshotAI: Kimi K2.7 Code
moonshotai/kimi-k2.7-code
$0.730$3.504.8×79%$1.42$0.982262K
Z.ai: GLM 5.1
z-ai/glm-5.1
$0.966$3.043.1×81%$1.48$1.15205K
Relace: Relace Search
relace/relace-search
$1.00$3.003.0×$1.50$1.18256K
Nous: Hermes 4 405B
nousresearch/hermes-4-405b
$1.00$3.003.0×$1.50$1.18131K
Qwen: Qwen3 Max Thinking
qwen/qwen3-max-thinking
$0.780$3.905.0×$1.56$1.06262K
Qwen: Qwen3 Max
qwen/qwen3-max
$0.780$3.905.0×80%$1.56$1.06262K
SpaceXAI: Grok 4.3
x-ai/grok-4.3
$1.25$2.502.0×84%$1.56$1.361M
SpaceXAI: Grok 4.20 Multi-Agent
x-ai/grok-4.20-multi-agent
$1.25$2.502.0×84%$1.56$1.362M
SpaceXAI: Grok 4.20
x-ai/grok-4.20
$1.25$2.502.0×84%$1.56$1.362M
OpenAI: GPT-3.5 Turbo Instruct
openai/gpt-3.5-turbo-instruct
$1.50$2.001.3×$1.63$1.554K
OpenAI GPT Mini Latest
~openai/gpt-mini-latest
$0.750$4.506.0×90%$1.69$1.09400K
OpenAI: GPT-5.4 Mini
openai/gpt-5.4-mini
$0.750$4.506.0×90%$1.69$1.09400K
Qwen: Qwen3 VL 235B A22B Thinking
qwen/qwen3-vl-235b-a22b-thinking
$0.980$3.954.0×$1.72$1.25131K
Thinking Machines: Inkling
thinkingmachines/inkling
$1.00$4.054.0×83%$1.76$1.281.0M
Z.ai: GLM 5V Turbo
z-ai/glm-5v-turbo
$1.20$4.003.3×80%$1.90$1.45203K
Z.ai: GLM 5 Turbo
z-ai/glm-5-turbo
$1.20$4.003.3×80%$1.90$1.45203K
OpenAI: o4 Mini High
openai/o4-mini-high
$1.10$4.404.0×75%$1.93$1.40200K
OpenAI: o4 Mini
openai/o4-mini
$1.10$4.404.0×75%$1.93$1.40200K
OpenAI: o3 Mini High
openai/o3-mini-high
$1.10$4.404.0×50%$1.93$1.40200K
OpenAI: o3 Mini
openai/o3-mini
$1.10$4.404.0×50%$1.93$1.40200K
Writer: Palmyra X5
writer/palmyra-x5
$0.600$6.0010.0×$1.95$1.091.0M
Meta: Muse Spark 1.1
meta/muse-spark-1.1
$1.25$4.253.4×88%$2.00$1.521.0M
Anthropic Claude Haiku Latest
~anthropic/claude-haiku-latest
$1.00$5.005.0×90%$2.00$1.36200K
Anthropic: Claude Haiku 4.5
anthropic/claude-haiku-4.5
$1.00$5.005.0×90%$2.00$1.36200K
Qwen: Qwen3.7 Max
qwen/qwen3.7-max
$1.48$4.423.0×80%$2.21$1.741M
OpenAI: GPT-5.6 Terra Pro
openai/gpt-5.6-terra-pro
$1.00$6.006.0×90%$2.25$1.451.1M
OpenAI: GPT-5.6 Terra
openai/gpt-5.6-terra
$1.00$6.006.0×90%$2.25$1.451.1M
Qwen: Qwen3.6 Max Preview
qwen/qwen3.6-max-preview
$1.03$6.166.0×$2.31$1.49262K
OpenAI: GPT-5 Image Mini
openai/gpt-5-image-mini
$2.50$2.000.8×90%$2.38$2.45400K
Qwen: Qwen3.8 Max
qwen/qwen3.8-max
$2.00$6.003.0×88%$3.00$2.361M
Google: Gemini 3.6 Flash
google/gemini-3.6-flash
$1.50$7.505.0×90%$3.00$2.051.0M
SpaceXAI: Grok 4.5
x-ai/grok-4.5
$2.00$6.003.0×85%$3.00$2.36500K
xAI: Grok Latest
~x-ai/grok-latest
$2.00$6.003.0×85%$3.00$2.36500K
Mistral: Mistral Medium 3.5
mistralai/mistral-medium-3-5
$1.50$7.505.0×$3.00$2.05262K
Google Gemini Flash Latest
~google/gemini-flash-latest
$1.50$7.505.0×90%$3.00$2.051.0M
Mistral Large 2407
mistralai/mistral-large-2407
$2.00$6.003.0×90%$3.00$2.36131K
Mistral: Mixtral 8x22B Instruct
mistralai/mixtral-8x22b-instruct
$2.00$6.003.0×90%$3.00$2.3666K
Mistral Large
mistralai/mistral-large
$2.00$6.003.0×90%$3.00$2.36128K
OpenAI: GPT-3.5 Turbo 16k
openai/gpt-3.5-turbo-16k
$3.00$4.001.3×$3.25$3.0916K
Google: Gemini 3.5 Flash
google/gemini-3.5-flash
$1.50$9.006.0×90%$3.38$2.181.0M
OpenAI: GPT-5.1-Codex-Max
openai/gpt-5.1-codex-max
$1.25$10.008.0×90%$3.44$2.05400K
OpenAI: GPT-5.1
openai/gpt-5.1
$1.25$10.008.0×90%$3.44$2.05400K
OpenAI: GPT-5.1-Codex
openai/gpt-5.1-codex
$1.25$10.008.0×90%$3.44$2.05400K
OpenAI: GPT-5
openai/gpt-5
$1.25$10.008.0×90%$3.44$2.05400K
Google: Gemini 2.5 Pro
google/gemini-2.5-pro
$1.25$10.008.0×90%$3.44$2.051.0M
Google: Gemini 2.5 Pro Preview 06-05
google/gemini-2.5-pro-preview
$1.25$10.008.0×90%$3.44$2.051.0M
Google: Gemini 2.5 Pro Preview 05-06
google/gemini-2.5-pro-preview-05-06
$1.25$10.008.0×90%$3.44$2.051.0M
AI21: Jamba Large 1.7
ai21/jamba-large-1.7
$2.00$8.004.0×$3.50$2.55256K
OpenAI: o3
openai/o3
$2.00$8.004.0×75%$3.50$2.55200K
OpenAI: GPT-4.1
openai/gpt-4.1
$2.00$8.004.0×75%$3.50$2.551.0M
Perplexity: Sonar Reasoning Pro
perplexity/sonar-reasoning-pro
$2.00$8.004.0×$3.50$2.55128K
Perplexity: Sonar Deep Research
perplexity/sonar-deep-research
$2.00$8.004.0×$3.50$2.55128K
Magnum v4 72B
anthracite-org/magnum-v4-72b
$3.00$5.001.7×$3.50$3.1816K
AionLabs: Aion-3.0
aion-labs/aion-3.0
$3.00$6.002.0×75%$3.75$3.27131K
Anthropic: Claude Sonnet 5
anthropic/claude-sonnet-5
$2.00$10.005.0×90%$4.00$2.731M
Anthropic Claude Sonnet Latest
~anthropic/claude-sonnet-latest
$2.00$10.005.0×90%$4.00$2.731M
OpenAI: GPT Audio
openai/gpt-audio
$2.50$10.004.0×$4.38$3.18128K
Cohere: Command A
cohere/command-a
$2.50$10.004.0×$4.38$3.18256K
OpenAI: GPT-4o (2024-11-20)
openai/gpt-4o-2024-11-20
$2.50$10.004.0×50%$4.38$3.18128K
Cohere: Command R+ (08-2024)
cohere/command-r-plus-08-2024
$2.50$10.004.0×$4.38$3.18128K
OpenAI: GPT-4o (2024-08-06)
openai/gpt-4o-2024-08-06
$2.50$10.004.0×50%$4.38$3.18128K
OpenAI: GPT-4o
openai/gpt-4o
$2.50$10.004.0×50%$4.38$3.18128K
Google: Nano Banana Pro (Gemini 3 Pro Image)
google/gemini-3-pro-image
$2.00$12.006.0×90%$4.50$2.91131K
Google Gemini Pro Latest
~google/gemini-pro-latest
$2.00$12.006.0×90%$4.50$2.911.0M
Google: Gemini 3.1 Pro Preview Custom Tools
google/gemini-3.1-pro-preview-customtools
$2.00$12.006.0×90%$4.50$2.911.0M
Google: Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
$2.00$12.006.0×90%$4.50$2.911.0M
Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
google/gemini-3-pro-image-preview
$2.00$12.006.0×90%$4.50$2.9166K
OpenAI: GPT-5.3 Chat
openai/gpt-5.3-chat
$1.75$14.008.0×90%$4.81$2.86128K
OpenAI: GPT-5.3-Codex
openai/gpt-5.3-codex
$1.75$14.008.0×90%$4.81$2.86400K
OpenAI: GPT-5.2-Codex
openai/gpt-5.2-codex
$1.75$14.008.0×90%$4.81$2.86400K
OpenAI: GPT-5.2 Chat
openai/gpt-5.2-chat
$1.75$14.008.0×90%$4.81$2.86128K
OpenAI: GPT-5.2
openai/gpt-5.2
$1.75$14.008.0×90%$4.81$2.86400K
Amazon: Nova Premier 1.0
amazon/nova-premier-v1
$2.50$12.505.0×75%$5.00$3.411M
OpenAI: GPT-5.4
openai/gpt-5.4
$2.50$15.006.0×90%$5.63$3.641.1M
MoonshotAI Kimi Latest
~moonshotai/kimi-latest
$2.90$14.004.8×90%$5.68$3.911.0M
MoonshotAI: Kimi K3
moonshotai/kimi-k3
$3.00$15.005.0×90%$6.00$4.091.0M
Anthropic: Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
$3.00$15.005.0×90%$6.00$4.091M
Perplexity: Sonar Pro Search
perplexity/sonar-pro-search
$3.00$15.005.0×$6.00$4.09200K
Anthropic: Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
$3.00$15.005.0×90%$6.00$4.091M
Anthropic: Claude Sonnet 4
anthropic/claude-sonnet-4
$3.00$15.005.0×90%$6.00$4.091M
Perplexity: Sonar Pro
perplexity/sonar-pro
$3.00$15.005.0×$6.00$4.09200K
OpenAI: GPT-4o (2024-05-13)
openai/gpt-4o-2024-05-13
$5.00$15.003.0×$7.50$5.91128K
OpenAI: GPT-5.4 Image 2
openai/gpt-5.4-image-2
$8.00$15.001.9×75%$9.75$8.64272K
Claude Opus 5
anthropic/claude-opus-5
$5.00$25.005.0×90%$10.00$6.821M
Anthropic: Claude Opus 4.8
anthropic/claude-opus-4.8
$5.00$25.005.0×90%$10.00$6.821M
Anthropic: Claude Opus Latest
~anthropic/claude-opus-latest
$5.00$25.005.0×90%$10.00$6.821M
Anthropic: Claude Opus 4.7
anthropic/claude-opus-4.7
$5.00$25.005.0×90%$10.00$6.821M
Anthropic: Claude Opus 4.6
anthropic/claude-opus-4.6
$5.00$25.005.0×90%$10.00$6.821M
Anthropic: Claude Opus 4.5
anthropic/claude-opus-4.5
$5.00$25.005.0×90%$10.00$6.82200K
OpenAI: GPT-5 Image
openai/gpt-5-image
$10.00$10.001.0×88%$10.00$10.00400K
OpenAI: GPT-5.6 Sol Pro
openai/gpt-5.6-sol-pro
$5.00$30.006.0×90%$11.25$7.271.1M
OpenAI: GPT-5.6 Sol
openai/gpt-5.6-sol
$5.00$30.006.0×90%$11.25$7.271.1M
Sakana: Fugu Ultra
sakana/fugu-ultra
$5.00$30.006.0×90%$11.25$7.271M
OpenAI: GPT Chat Latest
openai/gpt-chat-latest
$5.00$30.006.0×90%$11.25$7.27400K
OpenAI GPT Latest
~openai/gpt-latest
$5.00$30.006.0×90%$11.25$7.271.1M
OpenAI: GPT-5.5
openai/gpt-5.5
$5.00$30.006.0×90%$11.25$7.271.1M
OpenAI: GPT-4 Turbo
openai/gpt-4-turbo
$10.00$30.003.0×$15.00$11.82128K
OpenAI: GPT-4 Turbo Preview
openai/gpt-4-turbo-preview
$10.00$30.003.0×$15.00$11.82128K
Claude Opus 5 (Fast)
anthropic/claude-opus-5-fast
$10.00$50.005.0×90%$20.00$13.641M
Anthropic: Claude Fable Latest
~anthropic/claude-fable-latest
$10.00$50.005.0×90%$20.00$13.641M
Anthropic: Claude Fable 5
anthropic/claude-fable-5
$10.00$50.005.0×90%$20.00$13.641M
Anthropic: Claude Opus 4.8 (Fast)
anthropic/claude-opus-4.8-fast
$10.00$50.005.0×90%$20.00$13.641M
OpenAI: o1
openai/o1
$15.00$60.004.0×50%$26.25$19.09200K
Anthropic: Claude Opus 4.1
anthropic/claude-opus-4.1
$15.00$75.005.0×90%$30.00$20.45200K
Anthropic: Claude Opus 4
anthropic/claude-opus-4
$15.00$75.005.0×90%$30.00$20.45200K
OpenAI: o3 Pro
openai/o3-pro
$20.00$80.004.0×$35.00$25.45200K
OpenAI: GPT-4
openai/gpt-4
$30.00$60.002.0×$37.50$32.738K
OpenAI: GPT-5 Pro
openai/gpt-5-pro
$15.00$1208.0×$41.25$24.55400K
OpenAI: GPT-5.2 Pro
openai/gpt-5.2-pro
$21.00$1688.0×$57.75$34.36400K
Anthropic: Claude Opus 4.7 (Fast)
anthropic/claude-opus-4.7-fast
$30.00$1505.0×90%$60.00$40.911M
OpenAI: GPT-5.5 Pro
openai/gpt-5.5-pro
$30.00$1806.0×$67.50$43.641.1M
OpenAI: GPT-5.4 Pro
openai/gpt-5.4-pro
$30.00$1806.0×$67.50$43.641.1M
OpenAI: o1-pro
openai/o1-pro
$150$6004.0×$263$191200K

What this table is, and what it is not.

These are published list prices as the upstream catalogue reports them. They are not your negotiated rate, not your effective rate once caching and batch discounts land, and not a measurement of what any task costs you — token volume per task usually moves your bill more than unit price does. Use this to shortlist and to sanity-check an invoice. To find out what you are actually spending, you have to look at your own usage data. That is what the free teardown is for.