Google's full Gemini 3 Flash model, with a 1,048,576-token context window and $1.50/$9.00 per million token pricing for production-grade agent work.

Key specifications

Capability
89.0
Context window
1.0M
Max output
65.5K
Input $/1M
$1.50
Output $/1M
$9.00
License
Proprietary

Inputs and outputs

Recorded modality support shows what this model can accept and produce. An undocumented modality is shown as unknown rather than assumed to be unsupported.

Accepts

  • Text input
  • Image input
  • Video input
  • Audio input
  • PDF input

Produces

  • Text output

Capabilities

ReasoningFunction callingVisionAudioJSON mode

What is Gemini 3.5 Flash?

Overview

The gemini 3.5 flash model is the full, non-Lite Flash model in Google's Gemini 3 lineup — the version to reach for when a task needs more consistency than the Lite models offer but doesn't need a frontier-tier price tag. It's a fit for production agents, document-heavy RAG systems, and coding assistants where reliability on harder prompts matters more than shaving pennies off every call.

Token economics

The numbers here are the highest in this comparison: $1.50 per million input tokens and $9.00 per million output tokens, with cached input at $0.15 per million tokens. That output price is notably steeper than its own successor, Gemini 3.6 Flash, which charges $7.50 per million output tokens at the same input rate — worth noting if you're choosing between the two for a new build rather than maintaining an existing integration on 3.5 Flash.

A 1,048,576-token context window and 65,536-token output cap match the rest of the Gemini 3 family, so the pricing difference isn't buying extra room — it's paying for whatever changed under the hood between generations.

Modalities and reasoning

Input covers text, image, video, audio, and PDF; output is text only. Function calling and reasoning are both supported, which makes this a reasonable choice for agentic workflows that need to plan across multiple tool calls rather than just answer a single question.

  • Context window: 1,048,576 tokens
  • Max output: 65,536 tokens
  • Input: $1.50 per million tokens
  • Output: $9.00 per million tokens
  • Cached input: $0.15 per million tokens
  • Knowledge cutoff: January 2025

How it stacks up next to 3.1 Flash Lite and 3.6 Flash

Against Gemini 3.1 Flash Lite ($0.25 / $1.50), Gemini 3.5 Flash costs six times as much per input token and six times as much per output token — a jump that only makes sense if your task is actually hitting the Lite model's limits. Against Gemini 3.6 Flash, the newer full-size model in the family, 3.5 Flash charges the same for input but 20% more for output, which makes 3.6 Flash the more cost-efficient pick for new projects with similar quality needs.

Limitation worth knowing

The knowledge cutoff sits at January 2025, so for anything time-sensitive — recent events, current pricing elsewhere, newly released libraries — pairing this model with retrieval or search tools is safer than trusting it to know what happened since. Its capability score (89.0) is also a touch below the two newer models in this lineup (89.4 each), so it's worth testing against Gemini 3.6 Flash directly before committing to 3.5 Flash for a new integration.

How to evaluate Gemini 3.5 Flash

Use the recorded facts as a shortlist, then validate the model against representative inputs and production constraints.

Workload fit

Gemini 3.5 Flash is categorized for Audio-text-to-text, Automatic speech recognition, Image-text-to-text, Image-to-text, Text generation, Text-to-text AI models, Video-text-to-text. Its current record accepts text, image, video, audio, pdf and produces text. Confirm file formats, preprocessing, and provider-specific request schemas before implementation.

Capacity and cost

The directory records 1.0M context and 65.5K maximum output. Listed token prices are $1.50 input and $9.00 output per 1M tokens. Treat missing values as unknown and recheck current commercial terms.

Operational behavior

Recorded output speed is — and time to first token is —. Hosting route, region, prompt length, concurrency, and provider load can materially change both measurements.

Evidence boundary

This page separates sourced model facts from incomplete fields. Benchmark evidence is displayed only when its source is linked and verified. Before choosing Gemini 3.5 Flash, test task quality, tool reliability, safety behavior, data controls, rate limits, and total cost on the exact route you intend to use.

Strengths

  • + Registry-backed capabilities
  • + Provider: Google

Limitations

  • Draft record — editorial review required
  • Verify pricing and independent evaluations before approval

Frequently asked questions

Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens, with cached input priced at $0.15 per million tokens.

Related approved models in Automatic speech recognition. Compare specifications and verify fit for your workload.

Google

Gemini 3.5 Flash Lite

Google's low-cost Gemini 3 model with a 1,048,576-token context window, vision support, and $0.30/$2.50 per million token pricing for volume workloads.

1.0M contextCompare

Google

Gemini 3.6 Flash

Google's mid-tier Gemini 3 model: 1,048,576-token context, multimodal input, and $1.50/$7.50 per million token pricing for production agent workloads.

1.0M contextCompare

Google

Gemini 3.1 Flash Lite

Google's cheapest Gemini 3 model, with a 1,048,576-token context window and $0.25/$1.50 per million token pricing for high-volume, cost-sensitive work.

1.0M contextCompare

Google

Gemini Flash Latest

Google's fast, rolling-updated Gemini model with a 1M-token context window, reasoning, and tool calling at $1.50/$9.00 per million tokens.

1.0M contextCompare

Google

Gemini Flash-Lite Latest

Google's low-cost, high-volume Gemini model at $0.25/$1.50 per million tokens, with a 1M-token context window and reasoning support.

1.0M contextCompare