Open Gemma instruction model for efficient chat and self-hosted deployments

Key specifications

Capability
64.2
Context window
131.1K
Max output
8.2K
Input $/1M
$0.02
Output $/1M
$0.10
License
Open weights

Inputs and outputs

Recorded modality support shows what this model can accept and produce. An undocumented modality is shown as unknown rather than assumed to be unsupported.

Accepts

  • Text input
  • Image input
  • Audio input

Produces

  • Text output

Capabilities

ReasoningFunction callingVisionAudioJSON modeOpen weights

What is Gemma 4 E2B IT?

Editorial briefing: Gemma 4 E2B IT

Open Gemma instruction model for efficient chat and self-hosted deployments This is a source-backed review draft for Gemma 4 E2B IT from Google. It documents what the registry says and calls out what still needs a human editor to verify; it does not turn unverified fields into product claims.

What the source record says

The source record lists a release date of 2026-04-02 and a knowledge cutoff of not supplied in this registry record. It identifies the model family as gemma and describes open weights as available. These fields are useful for narrowing a shortlist, but the linked primary source remains the reference for any current release or policy change.

Interface and modality fit

Gemma 4 E2B IT accepts text, image, audio and returns text according to the registry. That makes the model a potential fit only where those interface assumptions match the product workflow. A team should test its own file types, tool schema, safety settings, and deployment constraints rather than assuming that a listed modality works identically in every provider surface.

Limits and operating assumptions

The listed context window is 131,072 tokens, with a listed maximum output of 8,192 tokens. The source declares reasoning, tool calling, file attachments, temperature control, structured output. Those flags describe an interface surface; they are not evidence of response quality, speed, tool reliability, or availability.

Recorded evaluation evidence

  • The registry record does not include a benchmark result for this model. That is an evidence gap, not a quality judgement.

A benchmark number is meaningful only alongside its task, harness, date, and source. It should not be treated as a universal ranking, and it should be refreshed before an editor uses it in a comparison or recommendation.

How to evaluate this model

Start with a small, representative test set: the prompts people actually send, the tools the product calls, the languages and files it receives, and the failure modes users care about. Compare output quality, structured-output reliability, latency, token use, safety behavior, and operational support under the same conditions for every candidate. Keep the source link, test date, region, and provider route next to any decision so the review can be repeated later.

Human review checklist

Before approval, an editor should verify the linked source, confirm whether a currently supported API or hosted route exists, add first-party pricing only when it is published, check benchmark provenance, and record any important limitations. If a fact cannot be verified, this page should say so plainly. That protects readers from stale model pages and gives search engines a materially useful, evidence-led resource rather than a generated catalogue entry.

How to evaluate Gemma 4 E2B IT

Use the recorded facts as a shortlist, then validate the model against representative inputs and production constraints.

Workload fit

Gemma 4 E2B IT is categorized for Audio-text-to-text, Automatic speech recognition, Image-text-to-text, Image-to-text, Text generation, Text-to-text AI models. Its current record accepts text, image, audio and produces text. Confirm file formats, preprocessing, and provider-specific request schemas before implementation.

Capacity and cost

The directory records 131.1K context and 8.2K maximum output. Listed token prices are $0.02 input and $0.10 output per 1M tokens. Treat missing values as unknown and recheck current commercial terms.

Operational behavior

Recorded output speed is — and time to first token is —. Hosting route, region, prompt length, concurrency, and provider load can materially change both measurements.

Evidence boundary

This page separates sourced model facts from incomplete fields. Benchmark evidence is displayed only when its source is linked and verified. Before choosing Gemma 4 E2B IT, test task quality, tool reliability, safety behavior, data controls, rate limits, and total cost on the exact route you intend to use.

Strengths

  • + Registry-backed capabilities
  • + Provider: Google

Limitations

  • Draft record — editorial review required
  • Verify pricing and independent evaluations before approval

Frequently asked questions

Gemma 4 E2B IT is a Google model listed in the provider’s public registry. Open Gemma instruction model for efficient chat and self-hosted deployments

Related approved models in Automatic speech recognition. Compare specifications and verify fit for your workload.

Google

Gemini 3.5 Flash Lite

Google's low-cost Gemini 3 model with a 1,048,576-token context window, vision support, and $0.30/$2.50 per million token pricing for volume workloads.

1.0M contextCompare

Google

Gemini 3.6 Flash

Google's mid-tier Gemini 3 model: 1,048,576-token context, multimodal input, and $1.50/$7.50 per million token pricing for production agent workloads.

1.0M contextCompare

Google

Gemini 3.1 Flash Lite

Google's cheapest Gemini 3 model, with a 1,048,576-token context window and $0.25/$1.50 per million token pricing for high-volume, cost-sensitive work.

1.0M contextCompare

Google

Gemini 3.5 Flash

Google's full Gemini 3 Flash model, with a 1,048,576-token context window and $1.50/$9.00 per million token pricing for production-grade agent work.

1.0M contextCompare

Google

Gemini Flash Latest

Google's fast, rolling-updated Gemini model with a 1M-token context window, reasoning, and tool calling at $1.50/$9.00 per million tokens.

1.0M contextCompare

Google

Gemini Flash-Lite Latest

Google's low-cost, high-volume Gemini model at $0.25/$1.50 per million tokens, with a 1M-token context window and reasoning support.

1.0M contextCompare