Google launches EmbeddingGemma 2, an open multimodal embedding model for devices

Google has released EmbeddingGemma 2, an open 740-million-parameter model that maps text, code, images, audio and video into a single space and runs entirely on a phone or a small board. It is licensed under Apache 2.0, has had no safety tuning, and Google says performance is not equal across the 100 languages it supports.

Google has released an open model that searches text, code, images, audio and video, and runs on a phone. It needs about 191MB of memory for text work on a Pixel 11 Pro, or about 567MB for all five. Nothing has to leave the device.

EmbeddingGemma 2 maps all of it into one 768-dimension space, Google said in a blog post. It has 740 million parameters, of which 270 million are needed for text alone, with a 170-million vision encoder and a 300-million audio encoder loaded only if wanted. The licence is Apache 2.0.

An embedding model turns content into numbers so software can find things by meaning rather than by matching words.

That is the step that normally happens in somebody else’s cloud.

We reported in April that Gemma 4 arrived as four open-weight models under the same licence, running on phones, Raspberry Pi boards and consumer graphics cards. EmbeddingGemma 2 is built on that architecture and shares its text tokeniser and audio encoder, so the two run together in less combined memory.

The practical result is a retrieval pipeline with no network in it.

Google made the same argument in August when it built an offline translator on an $80 Raspberry Pi, for handling sensitive conversations without sending audio anywhere. Keeping the data on the device avoids the transfer, which is where most European data protection questions begin.

The company is not doing this out of generosity. It is building the default.

Google says the Gemma family has passed a billion downloads, with the first EmbeddingGemma accounting for 20 million of them. The new model holds 8,192 tokens across every modality, four times the first version, which works out as 29 images, 58 video frames or five and a half minutes of audio.

Vectors can be cut from 768 dimensions to 128, which Google says reduces storage up to sixfold.

There are limits, and Google lists them. The model has had no safety tuning and no output moderation, with the mitigation applied to the training data instead. It supports more than 100 languages, and Google says performance may not be equal across them.

The EU has 24 official languages.

It is the kind of capability Europe keeps saying it wants to hold locally, running on hardware a European company already owns, under a licence anyone can use, and it was built in Mountain View.

Original source Google launches EmbeddingGemma 2, an open multimodal embedding model for devices

Back to home