Polarison

Gemini 2.5 Flash vs Gemini 3.1 Pro: pricing and benchmarks

Gemini 2.5 Flash is 85% cheaper on input tokens ($0.30 vs $2.00 per 1M). Gemini 2.5 Flash is 79% cheaper on output tokens ($2.50 vs $12.00 per 1M). On our document Q&A workload (50,000 questions a month), Gemini 2.5 Flash costs $195.00 a month versus $1,160 for Gemini 3.1 Pro, 83% less. Gemini 3.1 Pro scores 155.0 on the Epoch Capabilities Index vs 140.5 for Gemini 2.5 Flash.

Prices verified September 14, 2026.

Which should you choose?

Overall capability
Gemini 3.1 Pro

Gemini 3.1 Pro scores 155.0 on the Epoch Capabilities Index vs 140.5 for Gemini 2.5 Flash.

Price
Gemini 2.5 Flash

Gemini 2.5 Flash costs $0.85 per 1M tokens (3:1 input-to-output mix) vs $4.50 for Gemini 3.1 Pro, 81% less.

Best value
Too close to call

A trade-off: Gemini 3.1 Pro is 14.5 ECI points more capable but costs 5.3x as much as Gemini 2.5 Flash.

Pricing and specs

Gemini 2.5 FlashGemini 3.1 Pro
ProviderGoogleGoogle
API model IDgemini-2.5-flashgemini-3.1-pro-preview
Input / 1M tokens$0.30$2.00
Cached input / 1M tokens$0.03$0.20
Output / 1M tokens$2.50$12.00
Long-context rates$4.00 in / $18.00 out above 200K prompt tokens
Context window1.05M tokens1.05M tokens
Max output66K tokens66K tokens
Batch discount

Benchmarks

Gemini 2.5 FlashGemini 3.1 Pro
Epoch Capabilities Index140.5155.0
GPQA Diamond· Science reasoning92.6%
FrontierMath (Tiers 1–3)· Advanced math59.6%
SimpleQA Verified· Factual accuracy73.5%
ARC-AGI-2· Abstract reasoning77.1%
DeepSWE· Coding11.7%
APEX-Agents· Agentic work1.8%33.5%

Source: Epoch AI, best recorded result per model (CC BY 4.0). A dash means no published score. See all benchmarks

Monthly cost for real workloads

WorkloadGemini 2.5 FlashGemini 3.1 ProDifference
Support chatbot
100,000 replies a month, ~1,500 input tokens (system prompt + history) and ~400 output tokens each.
$145.00$780.00Gemini 2.5 Flash 81% cheaper
Document Q&A (RAG)
50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each.
$195.00$1,160Gemini 2.5 Flash 83% cheaper
Coding agent
10,000 agent steps a month, ~40,000 input tokens each (70% read from the prompt cache) and ~2,000 output tokens.
$94.40$536.00Gemini 2.5 Flash 82% cheaper

Estimates use standard list prices and ignore cache-write surcharges. Different tokenizers can count the same text differently, so test with your own prompts.

Gemini 2.5 Flash vs Gemini 3.1 Pro cost calculator

Standard-tier list prices. Long-context rates apply automatically where the provider publishes a threshold. Excludes cache-write surcharges, taxes, batch and volume discounts.

Gemini 2.5 Flash pricing notes

Google’s previous-generation Flash model.

  • Audio input costs $1.00 per 1M tokens; text, image and video cost $0.30.
  • Output price includes thinking tokens.
Google official pricing ↗

Gemini 3.1 Pro pricing notes

Google’s Gemini 3.1 Pro model, currently offered as a preview.

  • Prompts over 200K tokens are billed at $4 input / $18 output per 1M tokens.
  • No free tier on the Gemini API.
  • Output price includes thinking tokens.
Google official pricing ↗

Frequently asked questions

Is Gemini 2.5 Flash cheaper than Gemini 3.1 Pro?
Gemini 2.5 Flash is 85% cheaper on input tokens ($0.30 vs $2.00 per 1M). Gemini 2.5 Flash is 79% cheaper on output tokens ($2.50 vs $12.00 per 1M). On our document Q&A workload (50,000 questions a month), Gemini 2.5 Flash costs $195.00 a month versus $1,160 for Gemini 3.1 Pro, 83% less.
Which is more capable, Gemini 2.5 Flash or Gemini 3.1 Pro?
Gemini 3.1 Pro scores 155.0 on the Epoch Capabilities Index vs 140.5 for Gemini 2.5 Flash.
How much does Gemini 2.5 Flash cost per 1M tokens?
Gemini 2.5 Flash costs $0.30 per 1M input tokens and $2.50 per 1M output tokens, and $0.03 per 1M cached input tokens on Google’s standard tier.
How much does Gemini 3.1 Pro cost per 1M tokens?
Gemini 3.1 Pro costs $2.00 per 1M input tokens and $12.00 per 1M output tokens, and $0.20 per 1M cached input tokens on Google’s standard tier.
Which has the larger context window, Gemini 2.5 Flash or Gemini 3.1 Pro?
Both offer a 1.05M-token context window.

Related comparisons