Skip to content

AI model by Google

Gemma 4 for business

Google's open-weight model family under Apache 2.0, in sizes from phone-friendly to 31B, with image input and function calling.

  • Reviewed on September 24, 2026
  • Google

Capability tiersRelative, not benchmarks

Reasoning Medium
Coding Medium
Vision Medium
Speed High
Price level Lowest
Context size High
Tool calling Medium
Structured output Medium

Key facts about Gemma 4

Model ID at review
Gemma 4 (E2B, E4B, 12B, 26B A4B, 31B)
Provider
Google
Main category
Open weights
Open weights
Yes, Apache 2.0
Last reviewed
September 24, 2026

Fit

Where it fits and where it does not

Good at

  • Running on your own servers or devices.

  • Keeping data inside your own environment.

  • Image input on all sizes, audio on the smaller ones.

  • Function calling for simple agents.

  • Many languages.

Not the right choice for

  • The hardest reasoning tasks, where large hosted models lead.

  • Teams without the skills or budget to run models.

  • Very long contexts beyond 256K tokens.

  • Cases where a hosted API would be simpler and cheaper.

Use cases

Business use cases we would use it for

Private assistants

Internal assistants where data must stay on your own infrastructure. See RAG knowledge assistants.

On-device features

Small sizes for features that run on phones or laptops.

Private classification

Tagging and extraction on sensitive data without an outside API.

Our notes

When we would choose it

Gemma 4 is Google's current open-weight model family, released under the Apache 2.0 license. It comes in several sizes, E2B, E4B, 12B, 26B A4B and 31B, with context windows of 128K tokens on the smallest two and 256K on the larger ones. All sizes accept text and images, the smaller ones also accept audio, and Google lists native function calling and support for more than 140 languages.

Open weights mean you can download the model and run it where you like: your own servers, a cloud of your choice or even a phone. That gives full control over data and versions. The trade-off is that you run the infrastructure, handle scaling and apply updates yourself.

We would consider Gemma when data rules require that nothing leaves your environment, when an app needs to work offline, or when volumes are high enough that owning the hardware is cheaper than paying per request. For many small teams, a hosted model is still simpler.

The price level shown here reflects free weights, not the cost of hardware and people. Before deciding, compare the total cost and quality with a hosted option such as Gemini Flash-Lite. See open weights models.

A good first test is to run the 12B or 26B size on a single rented GPU against a sample of your real tasks, and compare the answers with a hosted model. That gives you quality and cost figures for your own work before you invest in hardware.

Before you commit

Things to check before you commit

  • Hardware

    Estimate GPU or CPU needs for your chosen size and traffic.

  • Operations

    Someone must run, update and monitor the model.

  • Quality gap

    Compare with a hosted model on the same test set.

  • License

    Gemma 4 uses Apache 2.0. Confirm it fits your use.

Alternatives

Models to compare it with

Llama 4 Scout

An open-weight Llama 4 model with image input and a very long context window, for teams that want to host a capable model themselves.

Mistral Small 4

Mistral's efficient open-weight model under Apache 2.0, one hybrid model for instructions, reasoning and coding, with tool calling and a 256K window.

Qwen3-Coder

Qwen's open-weight coding models under Apache 2.0, from the efficient Qwen3-Coder-Next to the large 480B model, with tool use and long context.

Keep exploring

FAQ

Questions people ask us

Have a question that is not here? Ask us directly.

Start a project

Not sure which model fits? Ask us to evaluate your use case.

Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.

Your privacy choices

We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.