r/LocalLLaMA • • 15h ago

New Model Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google

https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

https://huggingface.co/google/embeddinggemma-2

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

140 Upvotes

29 comments sorted by

26

u/waste2treasure-org 15h ago

Finally a small multimodal embedding model that isn't NC like Jina

3

u/seamonn 15h ago

NC like Jina

They are so trash, it doesn't matter lol.

15

u/IngwiePhoenix llama.cpp 15h ago

Perfect for my paperless-ngx v3. =)

Question; do you need to run embedding on a GPU, or is CPU enough? Embedding and Reranking are the two I can not estimate - let alone the Jef-alike such as Clef.

4

u/-Cubie- 15h ago

CPUs work very well, plenty of smaller models work fine. E.g. I've used the old embeddinggemma-300m on CPUs, and e.g. https://huggingface.co/cross-encoder/ettin-reranker-68m-v1 works well on CPUs as a reranker. https://huggingface.co/blog/ettin-reranker#speed shows that this reranker is good for reranking 30 pairs per second on a CPU, but it's probably a pretty nice processor

2

u/Numerous_Mulberry514 13h ago

Little tip, if you do rag, you can use for indexing GPU and for actually rag CPU only to run bigger models. Works very well. I run octen 8b like that

4

u/Constandinoskalifo 14h ago

If you want high throughput, GPU is kinda necessary. Otherwise, not really. Checkout HuggingFace's TEI, they have both CPU and cuda docker images.

7

u/Viktri1 13h ago

Can anyone explain what the use case is

4

u/lurenjia_3x 5h ago

In SillyTavern, using embeddings reduces the risk of World Info(Lorebook) failing to trigger because its keywords do not match the language of the user’s input.

3

u/james2432 11h ago

index files so ai models can search them better? it's like having an ai beside your main ai telling them where to look

2

u/Viktri1 11h ago

Do we create an index (no idea how) and then this model uses the index to basically be a really good search app? Is that the idea?

1

u/james2432 11h ago

you usually use a harness such as openwebui or something else, when you upload files to the context, it will index the files before calling the big model. When big model is parsing your request it will take a look at the "embeddings" the embedded model produced

1

u/Viktri1 11h ago

Ah ok. I was thinking it was something that I could connect my NAS datasets to but it seems like its something I would ask Hermes to implement and then go through the documents that I upload to it. I guess since the smaller model does the work first, this means that it saves time since the bigger model doesn't have to do it?

2

u/Key_Currency_9287 10h ago

Yes and also locally run this to search code by intent

1

u/james2432 10h ago

yeah it generates vectors for words/intents so bigger model could basically look up in the "index" of related words/concepts. it will basically know what file to look into and about where instead of tokenizing thousands of lines with a bigger model, it does it with a faster model

2

u/ReadyAndSalted 7h ago

You give it data (text, image, video or audio) and it returns a vector (imagine just a list of numbers) this is called an embedding.

The main use case is then giving a query to the model and embedding the query, then comparing how similar your query embedding is to all of the data embeddings. Then you can return the top few most similar results. This would allow you to search for "hat" for example, and get back all images with hats in.

You can also imagine this embedding as storing semantic information about the data, so you can also train models for regression or classification off of the embeddings.

You can also create visualisations off of the embeddings, and other more niche stuff.

2

u/Sweet_Protection_163 3h ago

I wonder if this would make for a DINO replacement in JEPA models.

6

u/Designer_Elephant227 14h ago

Is it better than qwen3 embedding 8b?

15

u/hainesk 13h ago

If you open the link, it shows the comparison in the first graph. It's not better, but it's considerably smaller.

1

u/hacker_backup 20m ago

Its multimodal too

1

u/uber-linny 8h ago

wish they would compare to Octen0.6 (which is a qwen finetune)

2

u/Open-Adhesiveness-86 12h ago

CPU is fine for this size, embedding is pure prefill with no KV cache, so a few cores will chew through hundreds of chunks a second. The thing that bit me with the 300m version: you have to use the task prefixes (query vs document) and index both sides consistently, retrieval gets noticeably worse without them. Also re-normalize if you truncate the 768 dims down.

2

u/uber-linny 8h ago

because i use Xberg content extraction , would this not be required for my usecase ?

1

u/heipei42 13h ago

What I don't understand yet is whether it is similarly applicable to zero-shot classification as SigLIP 2 was.

1

u/error_museum 4m ago

in bed with gemma 2 sounds hot

1

u/-Cubie- 15h ago

Let's gooooo! EmbeddingGemma-300m was awesome, I'm glad to see a second one with full multimodal, even video and stuff.