Polarison

The Cheapest AI APIs Right Now, Ranked

Every model from OpenAI, Anthropic, Google, xAI and DeepSeek ranked by blended price per token, plus the caveats that can flip the ranking.

Updated September 14, 2026

If you are building something high-volume — classification, extraction, summaries, a chatbot with thousands of daily users — the price per token matters more than the top of a leaderboard. We ranked every model we track by a blended price that weights input tokens three to one against output, which is close to how most chat and retrieval apps use tokens.

Every tracked model, cheapest first

#ModelProviderInputOutputBlended
1GPT-5.6 LunaOpenAI$0.20$1.20$0.45
2DeepSeek V4.1 FlashDeepSeek$0.30$1.20$0.525
3Gemini 3.5 Flash-LiteGoogle$0.30$2.50$0.85
4Gemini 2.5 FlashGoogle$0.30$2.50$0.85
5Gemini 3.8 FlashGoogle$0.75$3.75$1.50
6Grok 4.3xAI$1.25$2.50$1.563
7DeepSeek V4 ProDeepSeek$1.32$3.96$1.98
8Claude Haiku 4.5Anthropic$1.00$5.00$2.00
9Grok 4.6xAI$2.00$6.00$3.00
10Gemini 3.5 FlashGoogle$1.50$9.00$3.375
11Gemini 2.5 ProGoogle$1.25$10.00$3.438
12Claude Sonnet 5Anthropic$2.00$10.00$4.00
13GPT-5.6 TerraOpenAI$2.00$12.00$4.50
14Gemini 3.1 ProGoogle$2.00$12.00$4.50
15GPT-5.6 SolOpenAI$4.00$20.00$8.00
16Claude Opus 5Anthropic$5.00$25.00$10.00
17GPT-6 AstraOpenAI$10.00$50.00$20.00
18Claude Fable 5.1Anthropic$10.00$50.00$20.00

USD per 1M tokens, standard tier. Blended = (3 × input + output) ÷ 4.

The five cheapest right now

  1. GPT-5.6 Luna (OpenAI) — $0.20 input / $1.20 output per 1M tokens. The GPT-5.6 model optimized for cost-sensitive, high-volume workloads.
  2. DeepSeek V4.1 Flash (DeepSeek) — $0.30 input / $1.20 output per 1M tokens. DeepSeek’s low-cost model with thinking and non-thinking modes, tool calls and vision input.
  3. Gemini 3.5 Flash-Lite (Google) — $0.30 input / $2.50 output per 1M tokens. Google’s lowest-cost Gemini 3.5 model for high-volume work.
  4. Gemini 2.5 Flash (Google) — $0.30 input / $2.50 output per 1M tokens. Google’s previous-generation Flash model.
  5. Gemini 3.8 Flash (Google) — $0.75 input / $3.75 output per 1M tokens. Google’s newest Flash model, with thinking, tool use and a 1M-token context window.

Caveats before you switch

  • Off-peak pricing. DeepSeek halves its rates outside peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday to Friday). We list the peak price, so DeepSeek is cheaper than this table suggests for most of the week.
  • Introductory prices end. Gemini 3.8 Flash costs $0.75 input / $3.75 output through December 31, 2026, then doubles to $1.50 / $7.50. OpenAI’s GPT-5.6 Sol is on promotional pricing through at least November 21, 2026.
  • Your input-to-output ratio matters. Long generations favor models with cheap output tokens. Retrieval and classification favor cheap input. Re-rank with the calculator using your own numbers.
  • Tokenizers differ. The same paragraph can be a different number of tokens on different models, so identical per-token prices do not guarantee identical bills.
  • Cheap is not free if quality drops. If a budget model fails often enough that you retry or escalate to a bigger model, the savings shrink quickly. Measure success rate alongside cost.