Google's reasoning-first Gemini, built for problems that need a million tokens of context in one pass.

Key specifications

Capability
87.2
Context window
1.0M
Max output
65.5K
Input $/1M
$1.25
Output $/1M
$10.00
License
Proprietary

Inputs and outputs

Recorded modality support shows what this model can accept and produce. An undocumented modality is shown as unknown rather than assumed to be unsupported.

Accepts

  • Text input
  • Image input
  • Audio input
  • Video input
  • PDF input

Produces

  • Text output

Capabilities

ReasoningFunction callingVisionAudioJSON mode

What is Gemini 2.5 Pro?

Gemini 2.5 Pro is Google's reasoning model for problems that do not fit anywhere else. The number that defines it is the context window: 1,048,576 tokens, roughly 1,500 pages held in working memory at once. Hand it an entire repository, a full deposition, or a year of meeting transcripts, then ask a question whose answer depends on all of it.

That capacity is the reason to choose this model. It is also the reason to think first.

What it accepts

Gemini 2.5 Pro is natively multimodal on the way in. It reads:

  • Text — documents, source code, transcripts
  • Images — screenshots, diagrams, scanned pages
  • Audio — calls, interviews, lectures
  • Video — recorded sessions and screen captures
  • PDF — parsed directly, with no separate extraction step

Output is text only. If half your job is generating images or speech, this model does not do that half.

Where it earns its place

Reasoning is on by default, and that changes what the model is good at. It works through a problem in steps instead of answering from the first plausible pattern. The practical effect shows up in work where being wrong is expensive: reconciling documents that contradict each other, tracing a bug across files that never mention each other, planning a migration where order matters.

Function calling and JSON mode are both supported, so it fits into agent loops and structured pipelines without post-processing hacks to force the shape of a response.

Where it does not

Deliberation costs time. On high-volume, low-stakes calls — classification, tagging, short replies — you are paying latency for thinking the task never needed. Gemini 2.5 Flash carries the same million-token window and returns faster, which is usually the better trade at volume.

The 65,536-token output ceiling is generous, not infinite. Book-length generation still has to be chunked and stitched.

Specifications

Field Value
Provider Google
Context window 1,048,576 tokens
Maximum output 65,536 tokens
Input text, image, audio, video, PDF
Output text
Reasoning Yes
Function calling Yes
JSON mode Yes
Released 16 June 2025

Pricing, quotas and regional availability change frequently, so we deliberately do not reproduce them here. Confirm current figures against your own account before you commit to a budget.

Compare it against the rest of the AI model directory, or read how it differs from the faster sibling in Gemini 2.5 Flash.

How to evaluate Gemini 2.5 Pro

Use the recorded facts as a shortlist, then validate the model against representative inputs and production constraints.

Workload fit

Gemini 2.5 Pro is categorized for Audio-text-to-text, Automatic speech recognition, Image-text-to-text, Image-to-text, Text generation, Text-to-text AI models, Video-text-to-text. Its current record accepts text, image, audio, video, pdf and produces text. Confirm file formats, preprocessing, and provider-specific request schemas before implementation.

Capacity and cost

The directory records 1.0M context and 65.5K maximum output. Listed token prices are $1.25 input and $10.00 output per 1M tokens. Treat missing values as unknown and recheck current commercial terms.

Operational behavior

Recorded output speed is — and time to first token is —. Hosting route, region, prompt length, concurrency, and provider load can materially change both measurements.

Evidence boundary

This page separates sourced model facts from incomplete fields. Benchmark evidence is displayed only when its source is linked and verified. Before choosing Gemini 2.5 Pro, test task quality, tool reliability, safety behavior, data controls, rate limits, and total cost on the exact route you intend to use.

Strengths

  • + 1,048,576-token context window holds an entire codebase or case file at once
  • + Reads text, images, audio, video and PDFs natively — no separate extraction step
  • + Reasoning enabled by default, which suits multi-step work where errors are costly
  • + Function calling and JSON mode make it drop-in for agents and structured pipelines

Limitations

  • Reasoning adds latency you are paying for even on trivial calls
  • Text-only output — it cannot generate images or speech
  • 65,536-token output cap still forces chunking for book-length generation
  • Pricing and rate limits vary by region and change often; confirm before budgeting

Frequently asked questions

Gemini 2.5 Pro is a Google model listed in the provider’s public registry. Google's proven reasoning model for coding, math, and multimodal analysis

Related approved models in Automatic speech recognition. Compare specifications and verify fit for your workload.

Google

Gemini 3.5 Flash Lite

Google's low-cost Gemini 3 model with a 1,048,576-token context window, vision support, and $0.30/$2.50 per million token pricing for volume workloads.

1.0M contextCompare

Google

Gemini 3.6 Flash

Google's mid-tier Gemini 3 model: 1,048,576-token context, multimodal input, and $1.50/$7.50 per million token pricing for production agent workloads.

1.0M contextCompare

Google

Gemini 3.1 Flash Lite

Google's cheapest Gemini 3 model, with a 1,048,576-token context window and $0.25/$1.50 per million token pricing for high-volume, cost-sensitive work.

1.0M contextCompare

Google

Gemini 3.5 Flash

Google's full Gemini 3 Flash model, with a 1,048,576-token context window and $1.50/$9.00 per million token pricing for production-grade agent work.

1.0M contextCompare

Google

Gemini Flash Latest

Google's fast, rolling-updated Gemini model with a 1M-token context window, reasoning, and tool calling at $1.50/$9.00 per million tokens.

1.0M contextCompare

Google

Gemini Flash-Lite Latest

Google's low-cost, high-volume Gemini model at $0.25/$1.50 per million tokens, with a 1M-token context window and reasoning support.

1.0M contextCompare