---
title: Gemma
description: "Google's family of downloadable AI models for generation and specialized tasks, with capabilities and licenses that vary by release."
canonical_url: 'https://darkfactory.dev/glossary/gemma'
markdown_url: 'https://darkfactory.dev/glossary/gemma.md'
collection: glossary
date_published: '2026-10-07T00:00:00-04:00'
date_modified: '2026-10-07T00:00:00-04:00'
---

# Gemma


## Definition

Gemma is Google's family of downloadable AI models for local or hosted deployment. Core models generate text and, in supported releases, accept other input modalities. The wider family also includes models built for specific tasks such as embeddings and safety classification.

Google DeepMind and other Google teams develop Gemma using research and technology shared with Gemini. Gemma and Gemini are distinct families. A common technological lineage does not make their model IDs, weights, or supported interfaces interchangeable.

## Origin and attribution

Google introduced Gemma on February 21, 2024 with 2B and 7B models, each available in pretrained and instruction-tuned forms. Jeanine Banks and Tris Warkentin wrote the announcement. Google says the name comes from the Latin word for a precious stone.

The family has since expanded through numbered generations and specialized branches. Gemma 4 launched on April 2, 2026. Credit for these releases belongs to the Google teams that developed and published them.

## Family, versions, and capabilities

Specify the generation and variant when describing a deployment. The October 7, 2026 Gemma 4 overview lists E2B, E4B, 12B, 31B, and 26B A4B variants. They have different architectures and memory requirements: 31B is dense, while 26B A4B uses a mixture of experts.

Gemma 4 supports text and image input across these variants. Native audio support is specific to E2B, E4B, and 12B, and context capacities also vary. A family-level mention therefore cannot establish the inputs or context size supported by one downloaded artifact.

EmbeddingGemma and ShieldGemma have different task contracts from core generative Gemma. Their names identify specialized branches; they should not be treated as synonyms for every Gemma model.

## Licensing and limits

Licensing changes across releases. Google publishes Gemma 4 under Apache 2.0, while the Gemma 3 model card points to the custom Gemma Terms of Use. The label open weights describes access to parameters; the license for a particular release specifies its permissions and conditions.

Google's Gemma 4 model card records limitations in factual accuracy, ambiguous language, and difficult tasks. Downloading the weights does not remove those limitations. It lets the operator control deployment and tuning while taking responsibility for evaluating the resulting system.

## Operational significance

Budget for the exact weights, their precision, and the memory consumed by the context cache. A sparse model's active parameter count is not the amount of memory needed to load all its experts. Quantization can reduce the footprint, but the chosen artifact still needs evaluation on the intended tasks.

Keep a record of the release, variant, tokenizer, prompt format, and any tuning. Compare latency and task quality using the hardware and input distribution the application will actually use.

## Distinguish it from nearby terms

- An open-weight model makes parameters available. Gemma names a particular family with release-specific licenses.
- A small language model is a size and deployment category. Gemma includes several sizes and specialized architectures.
- A multimodal model accepts more than one kind of input or output. Supported modalities depend on the Gemma variant.
- Fine-tuning changes model parameters. Choosing a Gemma release alone does not adapt it to a particular organization.

## Check your understanding

A team replaces a Gemma 3 deployment with Gemma 4 26B A4B. Which license, input-capability, and memory assumptions should it check again, and why is the active parameter count insufficient for planning the deployment?

## Also called

Gemma model

## Related terms

- [Gemini](https://darkfactory.dev/glossary/gemini)
- [Open-weight model](https://darkfactory.dev/glossary/open-weight-model)
- [Foundation model](https://darkfactory.dev/glossary/foundation-model)
- [Small language model (SLM)](https://darkfactory.dev/glossary/small-language-model)
- [Multimodal model](https://darkfactory.dev/glossary/multimodal-model)
- [Mixture of experts (MoE)](https://darkfactory.dev/glossary/mixture-of-experts)
- [Fine-tuning](https://darkfactory.dev/glossary/fine-tuning)
- [Quantization](https://darkfactory.dev/glossary/quantization)
- [Context window](https://darkfactory.dev/glossary/context-window)
- [Evaluation (eval)](https://darkfactory.dev/glossary/evaluation)

## Related factory areas

- [Model selection, routing & budgets](https://darkfactory.dev/factory/model-routing-budgets)
- [Execution environments, identity & secrets](https://darkfactory.dev/factory/execution-environments)
- [Verification, evaluation & quality truth](https://darkfactory.dev/factory/verification)

## Evidence and further reading

- [Google: Introducing Gemma](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-open-models/)
- [Google: Gemma models overview](https://ai.google.dev/gemma/docs)
- [Google: Gemma 4 launch](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)
- [Google: Gemma 4 model card](https://ai.google.dev/gemma/docs/core/model_card_4)
- [Google: Gemma 4 overview](https://ai.google.dev/gemma/docs/core)
- [Google: Gemma 3 model card](https://ai.google.dev/gemma/docs/core/model_card_3)
- [Google: Gemma Terms of Use](https://ai.google.dev/gemma/terms)
