Google released EmbeddingGemma 2 on October 6: a 740M-parameter open embedding model that maps text, code, images, video and audio into one 768-dimensional vector space (the model card counts four modalities, with code inside text), built on the Gemma 4 architecture and tagged Apache 2.0. Google’s X post calls it “our first natively multimodal open model engineered for on-device embeddings.” Google on X Unlike Reflection’s Beam, this one shipped: the weights sit in one ungated Hugging Face repo, google/embeddinggemma-2, as a 1.49 GB bfloat16 safetensors file, and the same checkpoint loads as a 270M text-only model, 440M with the image encoder, 570M with the audio encoder, or the full 740M. EmbeddingGemma 2 model card Google DeepMind’s launch post gives it an 8,192-token context window, four times EmbeddingGemma 1’s, and says that quantized on a Pixel 11 Pro it “requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.” Google blog
Every quality score is Google’s own: no paper, an MTEB submission open and unaccepted at launch, and no independent quality benchmark on launch day. The license is declared in repo metadata with no LICENSE file at the root, and the card still points at the Gemma Prohibited Use Policy. The multilingual text score barely moved from version 1 (+0.21) and English went down; code is the only text score with a meaningful gain. What changed is modality coverage, context and license. For retrieval teams the trade is one model and one index for mixed media, paid for with a full re-embed and a quality floor nobody outside Google has measured. S5 Labs has not run the model.
What is downloadable, and what the license page says
One official repo exists, with gated: false. Hugging Face API EmbeddingGemma 1 was manually gated under the gemma license tag. Hugging Face API, EmbeddingGemma 1 The metadata reads license: apache-2.0, and the card header links a Gemma 4 license page that redirects to Google’s own copy of the Apache License 2.0; our researcher’s token-level comparison against the Apache Foundation text found no differences, a check a second reviewer did not repeat. Google Apache 2.0 page What the repo lacks is a LICENSE file: /resolve/main/LICENSE returns 404 and the 15-file listing has none, so the license exists as metadata and a link. The card also keeps a line from the old regime, “Deployments must adhere to the Gemma Prohibited Use Policy,” a document dated February 2024. We offer no legal reading of that line under Apache 2.0; we note only that the card still carries it. The draft MTEB metadata, written by a Google author, still lists license="gemma". That is draft text rather than a license statement, though it shows how recent the switch is. MTEB results PR #745 The family’s move off its custom license is in our Gemma 4 article; the open models guide separates headline licenses from their texts.
Day-one tooling was broad: transformers v5.19.0, published at 16:39 UTC that day, ships the EmbeddingGemma2 class; sentence-transformers 6.1.0 or newer is required; llama.cpp merged text, vision and audio support, and ggml-org, Unsloth, ONNX, MLX, Ollama and LiteRT-LM all had builds. llama.cpp PR #30054 vLLM support is merged to main but not released, and SGLang’s PR is open. vLLM PR #60254 Model Garden is “coming soon” per Google. Google blog
One checkpoint, four load options, one vector space
Video shares the 170M vision encoder, which is why the developer guide labels the 440M setup “Text, images, and video,” though the LiteRT-LM card lists only text and images for its 440M bundle. Developer guide The claim that matters for builders is the guide’s statement that “In sentence-transformers, all four encoder setups load from the same checkpoint, so they share one vector space”: a query embedded on the text-only model can be matched against documents embedded with the full model, and adding an encoder later does not mean recomputing what you have. That is Google’s statement, untested by us, and the card warns that encoder-omission loading “differs among model libraries.”
For retrieval builders the change is in pipeline shape. Google’s AI Edge post says the model reduces the latency and memory overhead of chaining separate captioning, speech-to-text and text-embedding models, and that its Video Moments Finder demo works without transcribing audio or generating captions. Google AI Edge post Our RAG pipeline guide compares text-only embedders and does not cover multimodal ones.
The 8,192 tokens are shared by every modality in one input: an image costs 280 tokens by default (about 29 per context), a video frame 140 (about 58 frames by arithmetic, though the shipped processor defaults to 1 fps and a max_frames of 32), and audio 25 tokens per second of mono 16 kHz, about 5.5 minutes. Processor config Interleaved inputs use <|image|>, <|video|> and <|audio|> placeholders and return one embedding; task prefixes apply to text only.
The hosted counterpart is Gemini Embedding 2, covered in our Gemini 3 Flash preview piece. It bills $12.00 per million video tokens on the standard paid tier ($6.00 on batch); EmbeddingGemma 2 has no per-call fee. Gemini API pricing One caution is ours: Google says embeddings from different modalities “can be compared on their semantic similarity” but publishes no calibration guidance, and cross-modal and same-modality cosine scores plausibly sit on different distributions. Until you have measured that, retrieve per modality and merge.
Two memory numbers from Google that do not reconcile
Disk sizes are unambiguous. The full bf16 checkpoint is 1.49 GB; weights only, our arithmetic gives 542 MB for the text-only load and 877 MB with vision. The LiteRT-LM bundles, quantization-aware trained with int4 transformer weights, int8 vision and mixed int2, int4 and int8 audio, are 165 MB for text, 388 MB for text plus vision and 485 MB for the full model. LiteRT-LM model card
Runtime memory is where Google’s figures diverge. The blog’s “~191MB” and “~567MB” are Pixel 11 Pro “active RAM” with quantization. The LiteRT-LM card measures the same phone at 112 MB for text and 127.5 MB with every encoder loaded, as CPU memory on the TPU backend with accelerator memory excluded, at 128 text tokens, a 70-token image and five to ten seconds of audio: short inputs, not the 8K maximum. Google does not reconcile the two, and the card says memory is “not directly comparable across operating systems,” so cite each to its source and do not blend them. Latency, Google-measured and not reproduced: the LiteRT-LM card lists 8.3 ms per text embedding on the Pixel 11 Pro TPU and 37.3 ms “Text + Vision latency” on a MacBook Pro M5 GPU, at the 70-token image setting rather than the 280-token default. LiteRT-LM model card The AI Edge post describes that 37.3 ms as visual embedding time, 26.9 images a second, on a “MacBook M5 Pro GPU”. Google AI Edge post Two notes from the main model card: run in bfloat16 or float32, because float16 returns NaN or silently degraded embeddings; and every benchmark uses the full-precision checkpoint, so no quality number exists for any int4 build. EmbeddingGemma 2 model card
The benchmarks are Google’s, and the general text scores barely moved
| Benchmark (768d, self-reported) | EmbeddingGemma 1 | EmbeddingGemma 2 |
|---|---|---|
| MTEB multilingual v2, Mean(Task) | 61.15 | 61.36 |
| MTEB English v2 | 69.67 | 68.46 |
| MTEB Code v1, NDCG@10 | 68.76 | 78.68 |
| MIEB lite, Mean(TaskType) | — | 64.64 |
| MMEB v2 overall (image 57.28, video 50.67, VisDoc 67.84) | — | 59.01 |
| MSEB retrieval, MRR@10 | — | 69.54 |
| MAEB, Mean(Task) | — | 49.39 |
Google’s MTEB submissions register these under modality-specific loads (text-only for the MTEB text rows, text-vision for MIEB, text-audio for MAEB), not one 740M entry.
Read as printed from each card (the two may differ in harness): multilingual is flat, English fell 1.21 points, and code rose 9.92, the gain the blog leads with. EmbeddingGemma 1 model card A team on EmbeddingGemma 1 for English text has no quality reason to move.
The sub-1B framing needs care. Google claims “leading scores among sub-1B multimodal embedders” across benchmarks including MTEB Code and MAEB, yet its own MAEB chart plots jina-embeddings-v5-omni-nano above EmbeddingGemma 2, and in the MTEB results repository that model scores 49.69 against EmbeddingGemma 2’s 49.39. It is licensed CC BY-NC 4.0, and its parameter count is under 1B by the Hugging Face API and about 1.04B by jina’s card. jina-embeddings-v5-omni-nano On MIEB lite, Google’s submitted Mean(Task) of 58.74 leads everything under 2B in that repository and beats BidirLM-Omni-2.5B and e5-omni-3B, but trails LCO-Embedding-Omni 3B and 7B and e5-omni-7B. MTEB results PR #745 On MMEB v2 the entry is self-reported, and the leaderboard’s own scoring over its score files puts the model 55th of the 65 entries that report image, video and visual-document scores (44th of the 45 that cover all 78 tasks), far below leaders at 80 and above; Qwen3-VL-Embedding-2B, Apache 2.0 and without audio, scores 73.25, self-reported, at 2.86 times the parameters. MMEB leaderboard
MTEB itself has not accepted the numbers. Asked for the instruction template, Google’s author wrote “I am unable to disclose the exact setup for evals,” adding that the instructions come from the provided prompts and that the results are reproducible with mteb through sentence-transformers. A maintainer then replied that the result “has to be computed using the reference implementation” and that, if the maintainers re-evaluate, “we can’t accept the current submitted scores.” mteb PR #5588 Our verifier recomputed 61.36, 78.68 and 49.39 from the PR’s own task files, so the arithmetic holds; the harness question is open. The accurate word is unaccepted, not disputed.
Migrating means re-embedding, and choosing a dimension
Matryoshka truncation to 512, 256 or 128 dimensions is where “up to 6x” storage comes from (768 divided by 128); the AI Edge post’s “up to 8x” has no supported dimension behind it. From the card’s own table, 512 dimensions keep 99.7% of the multilingual score, 98.2% of code and 98.9% of MMEB; 256 keeps 95% to 99% across the board; 128 keeps 90.8% of code but 77.4% of MMEB and 81.6% of MSEB. The card says 128 “is best suited to text-only workloads.” Truncated vectors must be re-normalized before cosine similarity, and skipping that step, the card says, “degrades ranking quality silently”; queries and documents must share a dimension. Qdrant, an early-access partner, reports that 1-bit TurboQuant on the full 768-dimensional vectors, with rescoring off, kept 99.0% of exact float32 nDCG@10 at 104 bytes per vector, averaged over five BEIR text sets, though only about 74% of top-10 results matched exact search. It is a third-party compression test, for text only. Qdrant
Vectors from version 1 do not carry over. Both models use 768 dimensions and the identical task: search result | query: and title: none | text: prompt strings, which makes mixing indexes tempting, but the weights differ and Google nowhere says the spaces are compatible. Our position: treat them as incompatible and re-embed. Chunking policy will then move storage more than dimension does, by our arithmetic: an hour of audio in 327-second windows is about 11 vectors; one vector per video frame is 3,600 an hour.
What to test before adopting it
- The license as you will rely on it: the Apache text at Google’s link, the Prohibited Use Policy line on the card, and whether a LICENSE file has appeared.
- Retrieval quality on your own data, per modality, with the exact build you will ship. Google’s scores are full precision; the int4 bundles have none.
- Cosine agreement between bf16-indexed documents and queries from a quantized build, and between text-only and full-model embeddings of the same input.
- Score calibration across modalities before using one threshold on a mixed corpus.
- Peak memory at your real input lengths, since Google measured at 128 text tokens and one 70-token image, and bfloat16 support on the target hardware.
- Whether MTEB accepts the submitted scores, and any independent quality evaluation; neither existed at launch.
Key Details
| Spec | Detail |
|---|---|
| Lab | Google DeepMind |
| Model ID | google/embeddinggemma-2 (one repo; four load options from one checkpoint) |
| Released | October 6, 2026 |
| Parameters | 740M full (744.37M); 270M text only; 440M text + image/video; 570M text + audio |
| Base | Gemma 4 architecture; shares Gemma 4’s tokenizer and audio encoder |
| Embedding dimensions | 768 native; Matryoshka 512 / 256 / 128 (128 best suited to text-only workloads per Google) |
| Context | 8,192 tokens shared across modalities; ~29 images, ~58 frames, ~5.5 min audio |
| Modalities | Text (incl. code), image, video, audio, interleaved |
| License | Apache 2.0 in repo metadata, ungated; no LICENSE file at repo root; card keeps Gemma Prohibited Use Policy line |
| Download | 1.49 GB bf16 safetensors; LiteRT-LM int4 bundles 165 / 388 / 485 MB; GGUF, ONNX, MLX, Ollama |
| Precision | bf16 or fp32 only; fp16 produces NaN or degraded output |
| Benchmarks | All self-reported, full precision; MTEB submission open and unaccepted; MMEB v2 59.01 self-reported |
| On-device memory | Google: ~191 MB text / ~567 MB full “active RAM” on Pixel 11 Pro (quantized); LiteRT card: 112 / 127.5 MB CPU memory at short inputs; not reconciled |
| Independent tests | No independent quality benchmark as of October 6; one third-party text-only compression test (Qdrant, early access) |
Sources
- EmbeddingGemma 2 announcement — Google on X
- EmbeddingGemma 2: an open, lightweight multimodal embedding model — Google
- EmbeddingGemma 2: the developer guide — Google for Developers
- Google AI Edge with EmbeddingGemma 2 — Google for Developers
- EmbeddingGemma 2 model card — Hugging Face
- google/embeddinggemma-2 repo metadata and file listing — Hugging Face API
- processor_config.json and config_sentence_transformers.json — Hugging Face
- Apache License 2.0 page linked from the card — Google AI for Developers
- embeddinggemma-2-740m-litert-lm model card — LiteRT Community
- EmbeddingGemma 1 model card — Hugging Face
- results: EmbeddingGemma 2 (PR #745) and model: EmbeddingGemma 2 (PR #5588) — embeddings-benchmark on GitHub
- MMEB Leaderboard score files — TIGER-Lab
- model: support embeddinggemma2 (PR #30054) — llama.cpp on GitHub
- Support EmbeddingGemma2 multimodal pooling (PR #60254) — vLLM on GitHub
- EmbeddingGemma 2 with 1-bit quantization — Qdrant
- jina-embeddings-v5-omni-nano model card — Jina AI
- Qwen3-VL-Embedding-2B model card — Alibaba Qwen
- Gemini API pricing — Google AI for Developers
