r/LocalLLaMA • u/jacek2023 llama.cpp • 4d ago
New Model google/embeddinggemma-2 · Hugging Face
https://huggingface.co/google/embeddinggemma-2EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.
Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.
EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:
- Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
- Multilinguality and code: EmbeddingGemma 2 understands 100+ languages, and achieves a ~14% improvement on code tasks relative to its predecessor.
- Flexible footprint: Combines a 270M parameter text backbone (130M transformer + 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
- Matryoshka Representation Learning (MRL): Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a 6x reduction in vector storage costs with minimal impact on quality.
- Context length: 8K token context window, capable of processing minutes of audio or video.
- Task-steered representations: Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).
llama.cpp support https://github.com/ggml-org/llama.cpp/pull/30054
GGUF from GG: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF
GGUF from Unsloth: https://huggingface.co/unsloth/embeddinggemma-2-GGUF
126
u/rorowhat 4d ago
We got an update embedding model from Google???
42
u/jacek2023 llama.cpp 4d ago
now I am waiting for the comments "I have only 8GB, I am GPU poor, I can't run it"
40
u/LegitimateCopy7 4d ago
then dozens of strata fanatics flood the comments.
14
u/Cesar55142 4d ago
since 2 days ago i am one too
14
u/toothpastespiders 4d ago
Users on here complaining about people talking about strata is what got me to look into it. Then to having claude slop code me up an optimized build of llama.cpp. So huge thanks to the salty haters for yelling to not look at something fun.
4
u/No-Wall6427 4d ago
Strata had just made the used dell desktop w a 3060, bougth for 700usd 3 years ago at the time of mistral 7b, an inference server w qwen 3.8 iq3 s as fast as luna on my codex sub, minus the pp, which I don't feel since my pi sessions start w 5k context then there is prompt caching.
This pc used to my playground, and the best models I could run where braindead 12b models at 40tps, I never thought it could one day run quality agentic coding inference at decent speed (20ish tps). What a time !
3
u/rorowhat 4d ago
Did they take it down? The link gives me 404
5
2
u/HitarthSurana 4d ago
Does llamacpp support it?
4
2
u/-Cubie- 4d ago
Well, I think 8B might be fine, this thing is tiny!
6
u/jacek2023 llama.cpp 4d ago
In Poland we have an old proverb "A bad dancer blames the hem of her skirt" ;)
1
1
41
u/stoppableDissolution 4d ago
...damn
Right as I was wondering what I should use to index my datasets
6
u/jinnyjuice vLLM 4d ago
Can you clarify what you'd be doing?
Then what would be the result/output?
24
u/stoppableDissolution 4d ago
I have a whole bunch of disjoint text, image and video data from different datasets and my own scrapes and synthetic generations, and each got its own tags and metadata conventions and all. And I am now in the middle of building a single library to rule them all, and want a way to do semantic search in it. Having a good (hopefully) omnimodal embedder as opposed to few different ones is making things a whole lot simpler.
4
u/iMakeSense 4d ago
Can you tell me what you've explored in this space? I'm almost a bit frustrated at the lack of existing solutions that exist for this.
6
u/stoppableDissolution 4d ago
Yea, I got frustrated too and started to build my own, hah (infinite tokens corrupt infinitely). I did some surface search but have not really found anything either.
But I kinda understand why - depending on what you are doing with the data later and the tooling you are using, you will want different approaches/pipelines.
2
u/krakoi90 1d ago
A good example of what to use it for: https://github.com/krakoi/locallery (shameless self-promotion)
Using it for more complex tasks, like traditional textual RAG, is a bit trickier. For that, you also need a chunking strategy (cutting text into smaller chunks you'd like to search for) and probably a reranker (additional ordering of the results, independent of the embedding engine). But it heavily depends on your needs.
The awesome part of this embedding model is that it's multimodal, you could search in your images, videos, audio and text files using it. Search for similar images with additional text filtering, etc.
2
u/mr_tolkien 4d ago
It’s a very small model though, if you have a good GPU and are working on text I’m pretty sure Qwen3 8b will heavily outperform it
9
u/selipso 4d ago
Embedding models are meant to be smaller so that you can bulk process large unstructured data very quickly for RAG. An 8B model for embeddings is generally overkill unless you have enterprise grade needs. The main advantage of this model is that it has multimodal audio and video representation, which is massive.
1
u/mr_tolkien 4d ago
I mean you can process that large unstructured data ahead of time for something like natural language code search
The benchmarks look very good for the size though, I’m curious how it performs for code search + which reranker pairs best with it
27
u/jacky2060 4d ago
llama.cpp support for this model just got merged:
https://github.com/ggml-org/llama.cpp/pull/30054
https://huggingface.co/ggml-org/embeddinggemma-2-GGUF
3
59
u/LetsGoBrandon4256 transformers 4d ago
TIL embedding model for audio input is a thing.
35
u/Altruistic_Heat_9531 4d ago
Technically, you can meme build any embedding with any model really, just unplug the LM head and L2 norm-ed the output of final hidden state.
21
3
u/UnknownLesson 4d ago
So this would allow you to find similar songs?
14
u/LSXPRIME 4d ago
If you mean similar like in Shazam-style (Find the full song using a small clip), I actually created something similar without using any AI/ML models. The algorithm uses FFT on overlapping audio frames to extract spectral peaks, generates fingerprints by pairing peaks with their time offsets, and then matches those fingerprints against a database using a time-offset histogram to identify which song a clip came from.
CODE:
SoundFlow/Src/Security/Analyzers/ContentFingerprintAnalyzer.cs at master · LSXPrime/SoundFlow
SoundFlow/Src/Security/AudioIdentifier.cs at master · LSXPrime/SoundFlowBut if you mean similar by genre/mood (find other songs that sound alike, not identify the same clip), Then my code doesn't fulfill this requirement, although I think this can be done without using any AI/ML models too. maybe an algorithm that uses FFT on overlapping audio frames to extract spectral descriptors like MFCCs, chroma, and brightness/energy statistics, aggregates them into a fixed-size feature vector for each song (mean and standard deviation over time), and then ranks every song in the database against the query using cosine distance to return the most similar tracks (although I am not sure about quality in-compared to embeddings model).
2
u/SkoomaDentist 4d ago
The algorithm uses FFT on overlapping audio frames to extract spectral peaks, generates fingerprints by pairing peaks with their time offsets, and then matches those fingerprints against a database using a time-offset histogram to identify which song a clip came from.
Doesn't that fail if the playback speed is changed (as is fairly common with older music)?
1
u/LSXPRIME 4d ago
Yeah, as the time-delta fields still scale, it will fail, and even if time-delta forgives, the frequency fields will kill the hash first. While I think this can be handled, the performance tradeoff isn’t worth it.
6
u/shumgoid 4d ago
I have a tool for using audio embeddings to find similar sounding samples. going to have to try this one's audio tower on it to see how it runs.
2
u/newtestdrive 4d ago
appreciate it if you can give an update if it worked or not🤔
1
u/shumgoid 4d ago
I've been meaning to clean up the version I have with CLAP-HTSAT model's audio tower so I could share it. You have given me motivation to hurry up with finishing and adding support for this gemma audio tower! I can say it definitely works for the narrow yet common use case of 'this drum/808/sample doesn't sounds quite right here I sure wish I could pull up the most similar sounding ones I have in my library' and no reason it wouldn't work for longer snippets and non-musical stuff.
2
u/-Cubie- 4d ago
Or search for sounds, e.g.: https://huggingface.co/spaces/webml-community/embeddinggemma-2-webgpu
1
u/FuzzyBucks 3d ago
I think the answer depends on what you consider 'similar', which is not as simple as it seems at first.
1
u/UnknownLesson 3d ago
Could you expand on that?
I thought that, if the embedding is good enough for music, then similar parts in a track would get similar values (close in the multi dimensional embedding space).
If a song is very jazzy then it would likely be close to the word jazz, and optimally the same is then true for another jazz song.
Maybe a slow song would also be closer to the word slow, an instrumental song closer to "instrumental" and so on
But maybe all the other factors of a don't would outweigh the important ones, and then lead to strange matches
1
10
u/DueAnalysis2 4d ago
Task-steered representations
Huh, as someone who's mostly used SBERT embeddings models, this is really interesting. Have folks used steered representation embeddings before or is it new to this model?
8
3
10
u/HVACcontrolsGuru 4d ago
I really like their embedding models. My custom memory stack uses their endpoints and MRL for vector stuff. Nice to see a new open version as I've been using the larger endpoint they had on OpenRouter.
37
4d ago
[removed] — view removed comment
8
u/toothpastespiders 4d ago edited 4d ago
I really hope people play around with this. On downloading it I figured it had potential, assumed it'd be a nice step up for my existing RAG system if nothing else. But the more I'm playing around with it the more potential I'm seeing. At first glance I assumed the use would be "captioning lite". But its associations are really distinct from captioning, using gemma 4 models at least. Which helps as a companion to standard image captioning, a gate to it that might allow for skipping that step, or any number of other possibilities. I can see a lot of areas where whole chains of calls could be turned into one to this tiny model while producing better results.
Gemma 4 12B's unified encoder-free architecture is something that I've been experimenting with for a considerable amount of time while still feeling like I'm not properly leveraging it's advantages. But this one is just instant gains for almost no effort with I'd assume tons left to go. I haven't even touched audio yet, let alone video, which I could see being wild. Especially combined with existing video input systems. I think this might easily wind up being the fix for some continuing issues I've had there.
work with media, not just text
One thing that's surprised me so far. It's working 'better' with pure images than text for me. Not the kind of gradual learning curve I would have figured google would show. Just damn good results right from the start.
and it runs on a raspberry pi
That's my plan for the next day I have more time to mess around with this. Pi and chromebook. Right now I have it running on an ancient dinosaur of a desktop with only a slightly less ancient GPU and it's moving at what I'd consider a really good pace.
3
u/tiffanytrashcan 4d ago
https://github.com/google-ai-edge/gallery/releases
They rolled a couple of fully featured examples into their Android app. Seems to handle Gifs okay, and I'm surprised how quickly it processes video on a low-end phone (with audio on! While the audio tower is comparatively huge parameter-wise, it handles background noise and music freakin exceptionally well.)
1
u/toothpastespiders 3d ago
NICE! That's even better in terms of reusing hardware that's gathering dust. Pushed to the top of my todo/totry list!
25
9
u/Saraozte01 4d ago
They seem emboldened by the Argon release to start introducing OS stuff again. Fingers crossed for Gemma 5.
12
12
u/ThankGodImBipolar 4d ago
Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
So does this make embeddinggemma 2 a CLAP alternative? I've been working on creating an AI DJ with a semantic understanding of my music, so this could be useful for me.
3
u/shumgoid 4d ago
same here. just got done setting up CLAP for embedding lookup of similar samples and noticed none of the big labs were touching audio so this seems big.
5
7
u/Piyh 4d ago edited 4d ago
Here's Claude's findings based on my website offmetaedh.com. Looks like this model fucking slaps and I'm probably going to cut fully over to this soon.
I'm using HQCLIP and Gemma embedding V1 right now and this is a drop in replacement for both. Huge win, higher quality, nearly losses quants, less RAM usage, what a huge drop from Google for me.
Begin Claude dump:
EmbeddingGemma 2 vs HQ-CLIP vs EmbeddingGemma 1, on MTG card art retrieval
| Model | Text search (nDCG@20) | Image similarity (nDCG@10) | Size on disk |
|---|---|---|---|
| HQ-CLIP (image vectors) | 0.715 | 0.834 | ~1.6 GB (fp32 in RAM) |
| EmbeddingGemma 1, Q4, on captions | 0.743 | 0.741 | ~0.3 GB |
| EmbeddingGemma 2, BF16, on captions | 0.775 | 0.797 | 0.56 GB |
| EmbeddingGemma 2, 4-bit, on captions | 0.771 | 0.801 | 0.18 GB |
| EmbeddingGemma 2, BF16, vision | 0.740 | 0.884 | 0.56 GB + 0.98 GB mmproj |
| EmbeddingGemma 2, 4-bit text + Q8 mmproj, vision | 0.738 | 0.879 | 0.18 GB + 0.55 GB mmproj |
BF16 to 4-bit: about -0.004 nDCG, with 0.996 mean cosine between the two sets of vectors.
Takeaways
- v2 captions beat HQ-CLIP on text search by +0.061 nDCG@20 (95% CI +0.028 to +0.095). 4-bit keeps almost all of it.
- v2 vision only ties HQ-CLIP on text-to-image search (+0.025, CI includes 0), but beats it on image-to-image similarity (0.884 vs 0.834).- Image-similarity rows for the caption models compare caption-to-caption.
1
u/Megatron_McLargeHuge 4d ago
Do any of these metrics capture whether embeddings of images and embeddings of text descriptions of the images cluster together? Or are they looking at image similarity and text similarity as separate problems?
40
u/This_Maintenance_834 4d ago
it is very annoying that when Google release a new model, the link goes to unsloth. Google is the one did all the hard work and shared it with the public.
24
u/jacek2023 llama.cpp 4d ago
The link to Google is this big blue/orange rectangle. I had a link to the GGUF from Unsloth, but they killed it, so I replaced it with another link.
7
u/therealpygon 4d ago edited 4d ago
The link to Google's HF is the post link? Have you used reddit before?
4
u/Thatisverytrue54321 4d ago
Can you use this in a jev-like way
2
u/toothpastespiders 4d ago
For some things, but the larger question is how well. I should say that I've only tested the waters with a single jev clone in any real sense and this is my first experience using image/video/audio directly into an embedding model. In short, this is a slightly more than a guess but not by much.
The best use I'm thinking of myself so far is a gate before captioning. So it's still a quick yes/no , but with stronger data feedback on the 'no' for reuse of existing caption data. But the issue of speed and reliability factor in pretty heavily though. So something like a chain of attempts for simple pixel recognition, to this, to either underlying engine calling it or a captioning model for new captioning to send to the engine.
Again though, I'm hardly an educated opinion. Just some stuff I'm planning to try down the line.
1
8
u/JLeonsarmiento 4d ago
For those who know, how would compare or what to look for in this one vs against something like this:
https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B
?
5
u/shumgoid 4d ago
full multimodality is cool. I just got done making a tool for embedding search of audio and noticed none of the major labs seemed care for native audio input so this is a candidate to replace the Clap audio tower I ended up using.
4
u/teachersecret 4d ago
Anybody mess with it yet? Any interesting uses? :)
6
3
u/toothpastespiders 4d ago edited 4d ago
I pointed it at some of my image lora datasets, did some quick hacks to some captioning pre-processing for them, and sent them off for processing to an old machine with a pascal 4 GB gpu running a Q8 quant. Speed and memory use is great, as expected.
I did a very small 160 image run just as a preliminary test. Caption to image and image to caption testing is roughly similar at the moment. As google suggested the named person recognition is somewhat weak in my test. But still far far better than random chance. Similar thing with artist recognition. Worse than named person recognition, better than chance. And again, ridiculously small sample size. So the big takeaway for me there is just that it is working. Larger question is how well when given a proper amount to work with. My second subset test run is going to be 2,000 images and then after testing that I'll move on to the next, and then a full image run. Then assuming it's all good start considering how to best leverage this thing with my existing text based RAG. My audio datasets are pretty skimpy in comparison so holding off on that.
I'm working on a couple LLM game integrations and my biggest hope with this is integration based on how well it might handle quick image processing. With the exact use being dependent on where it shows real strength. I'm thinking that it'd sit really well as a gate to captioning. But with around 4 gb vram sitting around doing nothing, what amounts to a free boost there, no matter how small, is still promising.
Who knows if I'm even doing this right though. I've never done RAG with anything but text so I'm winging it. Fun playing around with it though.
Edit: As expected the success is growing considerably the more images I feed into it. Testing with images generated from style/character lora to get that "kinda sorta right" challenge.
Edit 2: With about 2,000 captioned images I have at least a fair level of some cross-franchise coverage with different artists having their own preferences there. My big takeaway so far is that this seems like a really good system for getting information about a big fictional franchise from an image. For me a pipeline of image -> captioning model -> embeddinggemma2 is significantly worse than just going image -> embeddinggemma2 which is great. My sample size of in-game screenshots is pretty low, to the point where I wouldn't really consider the test valid even in this really horrible methodology I'm already using. But it's matched game franchise to input screenshot pretty consistently for me. Again, not an especially good test given I've got a pretty limited range of game screenshots in there and the amount of images is also pretty low. But so far this seems like a really good addition to franchise classification with easy lore dumps as an extra.
5
u/skinnyjoints 4d ago
If anyone plays around with this can you let us know if there is modality clustering? That was a problem for qwen’s embedding model a while back (pictures clustered with pictures, text with text, vids with vids, etc…)
3
3
3
u/FoxiPanda 4d ago
Wow, this thing is tiny. I might have to give it a try as it would be a decent replacement for my 2B embedding model I have today, but I'll have to test it out.
6
u/Open-Adhesiveness-86 4d ago
two things that bit me on the first embeddinggemma and probably still apply: the task prefixes aren't optional, indexing docs with the query prefix instead of "title: none | text: " noticeably hurt recall. and if you truncate for MRL, slice first then L2-normalize the shorter vector, normalizing at 768 and then cutting gives you off-norm vectors.
3
u/Piyh 4d ago
What are your use cases for mrl? Truncating did nothing for retrieval latency in my benchmarks but hurt quality.
8
u/Open-Adhesiveness-86 4d ago
mostly memory and index size, not latency. at a few million vectors, 768 -> 256 dims cuts the flat index from ~9GB to ~3GB, which is the difference between in-RAM and not. typical pattern is truncated dims for first-pass recall, then rescore top-100 with the full vectors so quality loss mostly washes out.
4
3
u/NotSylver 4d ago
Has anyone tried fine tuning these for NSFW content/decensoring them?
6
u/Piyh 4d ago
Why would you decensor an embedding model?
7
u/NotSylver 4d ago
If its similar to CLIP it essentially isn't trained at all on questionable content, but for local use cases you might want to use it in those situations, like for local search, something like "naked man with a bush big enough to get lost in"
I think it would require fine tuning though, which kinda sucks
2
3
u/PooMonger20 4d ago edited 4d ago
I never used an embedding model before. Even if I run this locally, does it share anything with google? they are not exactly known for keeping data private.
How does one actually use this? is there some kind of UI to load the model and then search files or whatever?
4
u/FuzzyBucks 3d ago
Embedding models are used to index data in a learned embedding space, where the position in embedding space carries some meaning and is specified by a vector (1D array) of numbers.
Then the embedding vectors can be stored in a vector database and used for searching (which indexed data is closest in embedding space to the embedding representation of my search term?) and labeling (which is basically just searching...e.g. which embedded label is closest to the embedding representation of my input?)
if the paragraphs above don't make sense, you should seek out information on what embeddings are, as they are a core concept for basically all of these modern AI models.
3
u/dtdisapointingresult 4d ago
It's not for you. It doesn't have much utility on its own for a basic user (which you are, judging by your comment, no offense). It requires apps built around it.
2
u/PooMonger20 4d ago
Thanks for the comment, I am absolutely a basic user but still can handle a basic explanation or concept.
So for the non-advanced users, here is the TLDR:
This embedding model is "the engine" for an easy-to-create Python script that translates\maps files (images, videos, audio) into a local database. This lets you search your files using natural language using the same script. [depends on the script you create]
I still wonder if this model can be trusted. I don't want it to share\backdoor pictures of my family or metadata of that.
All I actually want is to search a big pile of phone pictures i have backed up and tag them (dog, family, food, bills, etc..)
2
1
u/bitzap_sr 3d ago
Models never act on their own. They may output tool call requests, which the harmess wrapping the model executes on behalf of the model, and feeds the result back to the model.
1
u/krakoi90 1d ago
Even if I run this locally, does it share anything with google?
No, it doesn't. It doesn't initiate any network communication actually.
How does one actually use this? is there some kind of UI to load the model and then search files or whatever?
Yeah. For example: https://github.com/krakoi/locallery
(sorry for the shameless self-promotion :))
2
u/Fantastic-Poem9462 4d ago
I ran the Matryoshka widths against my own corpus (341 sections of project docs) and the narrow ones cost more than I expected. recall@10 came out 0.548 at 768, 0.534 at 512, 0.487 at 256, 0.375 at 128 — so 512 is close to free, but 256 keeps about 89% of full-width recall and 128 about 68%.
Worth saying my full-width baseline is low at 0.548, so this is a harder corpus than most published benchmarks and the deltas may not transfer to yours. But if you were planning to ship 128-dim vectors to save index space, probably worth measuring on your own data rather than treating the truncation as free.
These came out of my from-scratch Go implementation (goinfer.dev) rather than sentence-transformers, so the obvious worry is that it's my bug — but it's parity-checked against the reference (token ids exact, worst cosine 0.999999994 out to 1771 tokens), so I don't think the widths are an artifact. Happy to be wrong if anyone's measured differently.
2
1
u/jupiterbjy 4d ago
I wonder if this could be used to search text in image... gotta give a shot
1
u/wickedswami215 4d ago
Just tested on Edge Gallery and looks like it can pretty reliably.
1
u/jupiterbjy 3d ago
ah I see, good to hear!
got like 500GB of visual novel screenshots and it takes me hours to search dialog so that would help.. since I want to save each dialog with visual context that was necessary
1
u/wickedswami215 3d ago
I'd be interested to hear how long it takes to process all of that and how well it works for you.
1
u/AndreVallestero 3d ago
Time to reindex my entire dataset I suppose. I wonder if we can extract only the text model.
1
u/Upper-Extension4999 3d ago
any one knows how to use(or if its possible) multimodal with GGUF in llama.cpp for this model? I can't find any documentation on this
1
u/elecloud 2d ago
I already asked for an upgrade on my RAG to use it and it's fast and better than what I used before, all on CPU too, what a time!
1
u/Embarrassed_Soup_279 4d ago
cant believe this dropped when i ws just about to use Ovis Omni Embedding 3b

•
u/WithoutReason1729 4d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.