Google's fast, rolling-updated Gemini model with a 1M-token context window, reasoning, and tool calling at $1.50/$9.00 per million tokens.

Key specifications

Capability
89.0
Context window
1.0M
Max output
65.5K
Input $/1M
$1.50
Output $/1M
$9.00
License
Proprietary

Inputs and outputs

Recorded modality support shows what this model can accept and produce. An undocumented modality is shown as unknown rather than assumed to be unsupported.

Accepts

  • Text input
  • Image input
  • Video input
  • Audio input
  • PDF input

Produces

  • Text output

Capabilities

ReasoningFunction callingVisionAudioJSON mode

What is Gemini Flash Latest?

gemini flash latest is Google's fast Gemini tier that gets updated on a rolling basis instead of being frozen at a single snapshot — the "Latest" in the name means you're pointed at whatever the current fast build is, not a version you pin forever. It carries a 1,048,576-token context window and a 65,536-token output ceiling, with reasoning switched on by default so it can hold a chain of thought through multi-step tool calls rather than answering in one shot. At $1.50 per million input tokens and $9.00 per million output tokens, with cached input priced at $0.15 per million, it's built to be the model you call by default, not the one you reach for occasionally.

What "Latest" Means Here

Pointing at gemini flash latest instead of a dated snapshot buys convenience — Google rolls improvements into the alias without you changing a model string — at the cost of reproducibility. If an evaluation pipeline depends on stable outputs across weeks, a dated snapshot is the safer choice. If the goal is the current best fast model with no maintenance overhead, the alias does the job.

Context Window and Pricing in Practice

The quick numbers:

  • Context window: 1,048,576 tokens
  • Max output: 65,536 tokens
  • Input price: $1.50 per million tokens
  • Output price: $9.00 per million tokens
  • Cached input: $0.15 per million tokens
  • Input modalities: text, image, video, audio, pdf

1,048,576 tokens of context is enough to hold a large codebase, a long transcript, or a stack of PDFs in a single call, and the accepted input isn't limited to plain text. The economics matter more than the ceiling for most teams: at $1.50 in and $9.00 out per million tokens, a request with 50,000 input tokens and a 2,000-token reply costs roughly nine cents. Output is the expensive side of that ratio — six times the input rate — so prompts that ask for long, verbose answers cost more than prompts that ask for short, structured ones.

Reasoning and Tool Calling

Reasoning is on and function calling is supported, which makes this a workable base for agents that need to plan a sequence of tool calls rather than just answer a question. Vision is supported too, so image and video frames can sit in the same context alongside text without a separate pipeline.

Flash Latest vs Flash-Lite Latest

Google also ships Gemini Flash-Lite Latest at $0.25 input / $1.50 output per million tokens — roughly a sixth of Flash Latest's price. Both carry the same 1,048,576-token window and the same capability score of 89.0 in this listing, so the practical choice between them is about output quality per dollar on a specific workload, not raw ceiling. Flash Latest is the one to reach for when the extra cost buys noticeably better answers on real prompts; Flash-Lite Latest is the one to reach for when it doesn't.

One Thing to Watch

Output modality is text only — despite accepting image, video, audio, and pdf as input, it doesn't generate images or audio back. Teams building anything that needs generated visuals alongside a text answer will need a second model in the pipeline to cover that half of the job.

How to evaluate Gemini Flash Latest

Use the recorded facts as a shortlist, then validate the model against representative inputs and production constraints.

Workload fit

Gemini Flash Latest is categorized for Audio-text-to-text, Automatic speech recognition, Image-text-to-text, Image-to-text, Text generation, Text-to-text AI models, Video-text-to-text. Its current record accepts text, image, video, audio, pdf and produces text. Confirm file formats, preprocessing, and provider-specific request schemas before implementation.

Capacity and cost

The directory records 1.0M context and 65.5K maximum output. Listed token prices are $1.50 input and $9.00 output per 1M tokens. Treat missing values as unknown and recheck current commercial terms.

Operational behavior

Recorded output speed is — and time to first token is —. Hosting route, region, prompt length, concurrency, and provider load can materially change both measurements.

Evidence boundary

This page separates sourced model facts from incomplete fields. Benchmark evidence is displayed only when its source is linked and verified. Before choosing Gemini Flash Latest, test task quality, tool reliability, safety behavior, data controls, rate limits, and total cost on the exact route you intend to use.

Strengths

  • + Registry-backed capabilities
  • + Provider: Google

Limitations

  • Draft record — editorial review required
  • Verify pricing and independent evaluations before approval

Frequently asked questions

Gemini Flash Latest costs $1.50 per million input tokens and $9.00 per million output tokens, with cached input priced at $0.15 per million tokens.

Related approved models in Automatic speech recognition. Compare specifications and verify fit for your workload.

Google

Gemini 3.5 Flash Lite

Google's low-cost Gemini 3 model with a 1,048,576-token context window, vision support, and $0.30/$2.50 per million token pricing for volume workloads.

1.0M contextCompare

Google

Gemini 3.6 Flash

Google's mid-tier Gemini 3 model: 1,048,576-token context, multimodal input, and $1.50/$7.50 per million token pricing for production agent workloads.

1.0M contextCompare

Google

Gemini 3.1 Flash Lite

Google's cheapest Gemini 3 model, with a 1,048,576-token context window and $0.25/$1.50 per million token pricing for high-volume, cost-sensitive work.

1.0M contextCompare

Google

Gemini 3.5 Flash

Google's full Gemini 3 Flash model, with a 1,048,576-token context window and $1.50/$9.00 per million token pricing for production-grade agent work.

1.0M contextCompare

Google

Gemini Flash-Lite Latest

Google's low-cost, high-volume Gemini model at $0.25/$1.50 per million tokens, with a 1M-token context window and reasoning support.

1.0M contextCompare