HomeModelsGemini 1.5 Flash
GoogleFast Tier Vision

Gemini 1.5 Flash

Google · Released April 2024

Input Price

$0.07

per 1M tokens

Output Price

$0.30

per 1M tokens

Context Window

1.0M

tokens

Speed Score

99

out of 100

Benchmarks

Avg 81/100
Coding72/100
Reasoning76/100
Extraction93/100
Creative78/100
Vision85/100

Pricing

Input

Prompt / context tokens

$0.07/ 1M

Output

Generated tokens

$0.30/ 1M

Best For

extractionchatbotvisionanalysis

Gemini 1.5 Flash in context

On a typical 70% input / 30% output mix, Gemini 1.5 Flash blends to $0.14 per million tokens. That puts it above 3% of the 32 models tracked here, and 1 other models in the same budget tier undercut it. Tier is a capability label, not a price band — the spread inside one tier is often wider than the gap between tiers, which is why picking by tier alone tends to overspend.

Input runs $0.07 per million and output $0.30 — a 4.0x ratio. That ratio matters more than either number on its own: a summarisation or classification workload reads far more than it writes and will track the input price, while a code-generation or long-form writing workload inverts that and will track the output price. Compare models on the ratio your own traffic actually has, not on the input price alone.

The benchmark profile is uneven rather than flat: data extraction at 93/100 against coding at 72/100, a 21-point spread. An average hides that. If your workload sits on the strong end this model punches above its price; if it sits on the weak end, a cheaper model with a flatter profile will serve you better.

The context window is 1M tokens and the throughput score is 99/100. Context only earns its keep if you fill it, and filling it is also what makes a request expensive — a large window is an option, not a discount. For a budget-tier model, judge the speed score against what you are doing: it decides everything for interactive chat and almost nothing for overnight batch work.

Nothing cheaper in the table matches Gemini 1.5 Flash on combined coding and reasoning, so its $0.14 blended price is buying capability you cannot get for less right now. That is the case for paying it — and it is worth re-checking, because this is exactly the position that a new release takes away.

The closest comparison is GPT-5 Nano from OpenAI, at $0.11 per million against $0.14. 8 models in this tier score higher on combined coding and reasoning, which is the useful framing: the question is rarely whether a model is good, it is whether it is the cheapest thing that is good enough for the specific work you are sending it.

What it costs in production

At 10M + 2M tokens a month — a realistic mid-size production workload — Gemini 1.5 Flash runs about $1.35. There is no batch discount on this model, so that figure is the floor — the usual lever of shifting background work to a cheaper asynchronous tier is not available here.

Price History

DateInput / 1MOutput / 1MChange
Apr 2024$0.35$1.05Launch
Sep 2024Current$0.07$0.3079% cut

Strengths

  • Cheapest model with strong vision capabilities
  • 1M token context at minimal cost
  • Very fast response times
  • Good structured output

Avoid For

  • Advanced code generation
  • Complex multi-step reasoning