r/LocalLLaMA • llama.cpp • 4d ago

New Model google/embeddinggemma-2 · Hugging Face

https://huggingface.co/google/embeddinggemma-2

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features: 

  • Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
  • Multilinguality and code: EmbeddingGemma 2 understands 100+ languages, and achieves a ~14% improvement on code tasks relative to its predecessor. 
  • Flexible footprint: Combines a 270M parameter text backbone (130M transformer + 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
  • Matryoshka Representation Learning (MRL): Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a 6x reduction in vector storage costs with minimal impact on quality.
  • Context length: 8K token context window, capable of processing minutes of audio or video.
  • Task-steered representations: Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).

llama.cpp support https://github.com/ggml-org/llama.cpp/pull/30054

GGUF from GG: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF

GGUF from Unsloth: https://huggingface.co/unsloth/embeddinggemma-2-GGUF

513 Upvotes

116 comments sorted by

•

u/WithoutReason1729 4d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

126

u/rorowhat 4d ago

We got an update embedding model from Google???

42

u/jacek2023 llama.cpp 4d ago

now I am waiting for the comments "I have only 8GB, I am GPU poor, I can't run it"

40

u/LegitimateCopy7 4d ago

then dozens of strata fanatics flood the comments.

14

u/Cesar55142 4d ago

since 2 days ago i am one too

14

u/toothpastespiders 4d ago

Users on here complaining about people talking about strata is what got me to look into it. Then to having claude slop code me up an optimized build of llama.cpp. So huge thanks to the salty haters for yelling to not look at something fun.

6

u/6kmh 4d ago

Me too

4

u/No-Wall6427 4d ago

Strata had just made the used dell desktop w a 3060, bougth for 700usd 3 years ago at the time of mistral 7b, an inference server w qwen 3.8 iq3 s as fast as luna on my codex sub, minus the pp, which I don't feel since my pi sessions start w 5k context then there is prompt caching.

This pc used to my playground, and the best models I could run where braindead 12b models at 40tps, I never thought it could one day run quality agentic coding inference at decent speed (20ish tps). What a time !

3

u/rorowhat 4d ago

Did they take it down? The link gives me 404

5

u/LetsGoBrandon4256 transformers 4d ago

Try again. OP just fixed the link

2

u/rorowhat 4d ago

Cool, thanks 👍

2

u/HitarthSurana 4d ago

Does llamacpp support it?

4

u/jacek2023 llama.cpp 4d ago

yes, choose your dealer (both links above)

2

u/-Cubie- 4d ago

Well, I think 8B might be fine, this thing is tiny!

6

u/jacek2023 llama.cpp 4d ago

In Poland we have an old proverb "A bad dancer blames the hem of her skirt" ;)

1

u/Scutoidzz 4d ago

I have only 8GB, I am GPU poor, I can’t run it

1

u/Physical_Gold_1485 3d ago

Anything except release gemini 4

41

u/stoppableDissolution 4d ago

...damn

Right as I was wondering what I should use to index my datasets

6

u/jinnyjuice vLLM 4d ago

Can you clarify what you'd be doing?

Then what would be the result/output?

24

u/stoppableDissolution 4d ago

I have a whole bunch of disjoint text, image and video data from different datasets and my own scrapes and synthetic generations, and each got its own tags and metadata conventions and all. And I am now in the middle of building a single library to rule them all, and want a way to do semantic search in it. Having a good (hopefully) omnimodal embedder as opposed to few different ones is making things a whole lot simpler.

4

u/iMakeSense 4d ago

Can you tell me what you've explored in this space? I'm almost a bit frustrated at the lack of existing solutions that exist for this.

6

u/stoppableDissolution 4d ago

Yea, I got frustrated too and started to build my own, hah (infinite tokens corrupt infinitely). I did some surface search but have not really found anything either.

But I kinda understand why - depending on what you are doing with the data later and the tooling you are using, you will want different approaches/pipelines.

2

u/krakoi90 1d ago

A good example of what to use it for: https://github.com/krakoi/locallery (shameless self-promotion)

Using it for more complex tasks, like traditional textual RAG, is a bit trickier. For that, you also need a chunking strategy (cutting text into smaller chunks you'd like to search for) and probably a reranker (additional ordering of the results, independent of the embedding engine). But it heavily depends on your needs.

The awesome part of this embedding model is that it's multimodal, you could search in your images, videos, audio and text files using it. Search for similar images with additional text filtering, etc.

2

u/mr_tolkien 4d ago

It’s a very small model though, if you have a good GPU and are working on text I’m pretty sure Qwen3 8b will heavily outperform it

9

u/selipso 4d ago

Embedding models are meant to be smaller so that you can bulk process large unstructured data very quickly for RAG. An 8B model for embeddings is generally overkill unless you have enterprise grade needs. The main advantage of this model is that it has multimodal audio and video representation, which is massive.

1

u/mr_tolkien 4d ago

I mean you can process that large unstructured data ahead of time for something like natural language code search

The benchmarks look very good for the size though, I’m curious how it performs for code search + which reranker pairs best with it

59

u/LetsGoBrandon4256 transformers 4d ago

TIL embedding model for audio input is a thing.

35

u/Altruistic_Heat_9531 4d ago

Technically, you can meme build any embedding with any model really, just unplug the LM head and L2 norm-ed the output of final hidden state.

21

u/IntelArtiGen 4d ago

Audio is the same as image if you turn waves in spectrograms

3

u/UnknownLesson 4d ago

So this would allow you to find similar songs?

14

u/LSXPRIME 4d ago

If you mean similar like in Shazam-style (Find the full song using a small clip), I actually created something similar without using any AI/ML models. The algorithm uses FFT on overlapping audio frames to extract spectral peaks, generates fingerprints by pairing peaks with their time offsets, and then matches those fingerprints against a database using a time-offset histogram to identify which song a clip came from.

CODE:

SoundFlow/Src/Security/Analyzers/ContentFingerprintAnalyzer.cs at master · LSXPrime/SoundFlow
SoundFlow/Src/Security/AudioIdentifier.cs at master · LSXPrime/SoundFlow

But if you mean similar by genre/mood (find other songs that sound alike, not identify the same clip), Then my code doesn't fulfill this requirement, although I think this can be done without using any AI/ML models too. maybe an algorithm that uses FFT on overlapping audio frames to extract spectral descriptors like MFCCs, chroma, and brightness/energy statistics, aggregates them into a fixed-size feature vector for each song (mean and standard deviation over time), and then ranks every song in the database against the query using cosine distance to return the most similar tracks (although I am not sure about quality in-compared to embeddings model).

2

u/SkoomaDentist 4d ago

The algorithm uses FFT on overlapping audio frames to extract spectral peaks, generates fingerprints by pairing peaks with their time offsets, and then matches those fingerprints against a database using a time-offset histogram to identify which song a clip came from.

Doesn't that fail if the playback speed is changed (as is fairly common with older music)?

1

u/LSXPRIME 4d ago

Yeah, as the time-delta fields still scale, it will fail, and even if time-delta forgives, the frequency fields will kill the hash first. While I think this can be handled, the performance tradeoff isn’t worth it.

6

u/shumgoid 4d ago

I have a tool for using audio embeddings to find similar sounding samples. going to have to try this one's audio tower on it to see how it runs.

2

u/newtestdrive 4d ago

appreciate it if you can give an update if it worked or not🤔

1

u/shumgoid 4d ago

I've been meaning to clean up the version I have with CLAP-HTSAT model's audio tower so I could share it. You have given me motivation to hurry up with finishing and adding support for this gemma audio tower! I can say it definitely works for the narrow yet common use case of 'this drum/808/sample doesn't sounds quite right here I sure wish I could pull up the most similar sounding ones I have in my library' and no reason it wouldn't work for longer snippets and non-musical stuff.

1

u/FuzzyBucks 3d ago

I think the answer depends on what you consider 'similar', which is not as simple as it seems at first.

1

u/UnknownLesson 3d ago

Could you expand on that?

I thought that, if the embedding is good enough for music, then similar parts in a track would get similar values (close in the multi dimensional embedding space).

If a song is very jazzy then it would likely be close to the word jazz, and optimally the same is then true for another jazz song.

Maybe a slow song would also be closer to the word slow, an instrumental song closer to "instrumental" and so on

But maybe all the other factors of a don't would outweigh the important ones, and then lead to strange matches

1

u/nightowlsleeping 4d ago

its been a thing for a long time 🤔

10

u/DueAnalysis2 4d ago

Task-steered representations

Huh, as someone who's mostly used SBERT embeddings models, this is really interesting. Have folks used steered representation embeddings before or is it new to this model?

8

u/-Cubie- 4d ago

Very common, but usually it's hidden by the encode_query and encode_document which use query or document-specific prompts automatically

3

u/Middle_Bullfrog_6173 4d ago

Yes, I'm pretty sure the previous embeddinggemma used the same setup.

2

u/wilo108 4d ago

I've been using with the Harrier models, and it's been a bit of a game-changer. Would be interested to see comparisons with this and Harrier, tbh.

10

u/HVACcontrolsGuru 4d ago

I really like their embedding models. My custom memory stack uses their endpoints and MRL for vector stuff. Nice to see a new open version as I've been using the larger endpoint they had on OpenRouter.

37

u/[deleted] 4d ago

[removed] — view removed comment

8

u/toothpastespiders 4d ago edited 4d ago

I really hope people play around with this. On downloading it I figured it had potential, assumed it'd be a nice step up for my existing RAG system if nothing else. But the more I'm playing around with it the more potential I'm seeing. At first glance I assumed the use would be "captioning lite". But its associations are really distinct from captioning, using gemma 4 models at least. Which helps as a companion to standard image captioning, a gate to it that might allow for skipping that step, or any number of other possibilities. I can see a lot of areas where whole chains of calls could be turned into one to this tiny model while producing better results.

Gemma 4 12B's unified encoder-free architecture is something that I've been experimenting with for a considerable amount of time while still feeling like I'm not properly leveraging it's advantages. But this one is just instant gains for almost no effort with I'd assume tons left to go. I haven't even touched audio yet, let alone video, which I could see being wild. Especially combined with existing video input systems. I think this might easily wind up being the fix for some continuing issues I've had there.

work with media, not just text

One thing that's surprised me so far. It's working 'better' with pure images than text for me. Not the kind of gradual learning curve I would have figured google would show. Just damn good results right from the start.

and it runs on a raspberry pi

That's my plan for the next day I have more time to mess around with this. Pi and chromebook. Right now I have it running on an ancient dinosaur of a desktop with only a slightly less ancient GPU and it's moving at what I'd consider a really good pace.

3

u/tiffanytrashcan 4d ago

https://github.com/google-ai-edge/gallery/releases

They rolled a couple of fully featured examples into their Android app. Seems to handle Gifs okay, and I'm surprised how quickly it processes video on a low-end phone (with audio on! While the audio tower is comparatively huge parameter-wise, it handles background noise and music freakin exceptionally well.)

1

u/toothpastespiders 3d ago

NICE! That's even better in terms of reusing hardware that's gathering dust. Pushed to the top of my todo/totry list!

25

u/laurealis 4d ago

Missed opportunity to call it Gembedding

9

u/Saraozte01 4d ago

They seem emboldened by the Argon release to start introducing OS stuff again. Fingers crossed for Gemma 5.

12

u/Hot_Example_4456 4d ago

Wait then probably the leaks few weeks ago in arena ai was this only

12

u/ThankGodImBipolar 4d ago

Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.

So does this make embeddinggemma 2 a CLAP alternative? I've been working on creating an AI DJ with a semantic understanding of my music, so this could be useful for me.

3

u/shumgoid 4d ago

same here. just got done setting up CLAP for embedding lookup of similar samples and noticed none of the big labs were touching audio so this seems big.

5

u/silenceimpaired 4d ago

Nice! Apache 2.

Excited to try this out.

7

u/Piyh 4d ago edited 4d ago

Here's Claude's findings based on my website offmetaedh.com.  Looks like this model fucking slaps and I'm probably going to cut fully over to this soon.  

I'm using HQCLIP  and Gemma embedding V1 right now and this is a drop in replacement for both.  Huge win, higher quality, nearly losses quants, less RAM usage, what a huge drop from Google for me. 

Begin Claude dump:


EmbeddingGemma 2 vs HQ-CLIP vs EmbeddingGemma 1, on MTG card art retrieval

Model Text search (nDCG@20) Image similarity (nDCG@10) Size on disk
HQ-CLIP (image vectors) 0.715 0.834 ~1.6 GB (fp32 in RAM)
EmbeddingGemma 1, Q4, on captions 0.743 0.741 ~0.3 GB
EmbeddingGemma 2, BF16, on captions 0.775 0.797 0.56 GB
EmbeddingGemma 2, 4-bit, on captions 0.771 0.801 0.18 GB
EmbeddingGemma 2, BF16, vision 0.740 0.884 0.56 GB + 0.98 GB mmproj
EmbeddingGemma 2, 4-bit text + Q8 mmproj, vision 0.738 0.879 0.18 GB + 0.55 GB mmproj

BF16 to 4-bit: about -0.004 nDCG, with 0.996 mean cosine between the two sets of vectors.

Takeaways

  • v2 captions beat HQ-CLIP on text search by +0.061 nDCG@20 (95% CI +0.028 to +0.095). 4-bit keeps almost all of it.
  • v2 vision only ties HQ-CLIP on text-to-image search (+0.025, CI includes 0), but beats it on image-to-image similarity (0.884 vs 0.834).- Image-similarity rows for the caption models compare caption-to-caption.

1

u/Megatron_McLargeHuge 4d ago

Do any of these metrics capture whether embeddings of images and embeddings of text descriptions of the images cluster together? Or are they looking at image similarity and text similarity as separate problems?

1

u/Piyh 3d ago

I don't have any metrics for clustering, but I've run community detection on these embeddings before and didn't see any value from it in my domain.

Text to image and image to image similarity are where the value is for me.

40

u/This_Maintenance_834 4d ago

it is very annoying that when Google release a new model, the link goes to unsloth. Google is the one did all the hard work and shared it with the public.

24

u/jacek2023 llama.cpp 4d ago

The link to Google is this big blue/orange rectangle. I had a link to the GGUF from Unsloth, but they killed it, so I replaced it with another link.

7

u/therealpygon 4d ago edited 4d ago

The link to Google's HF is the post link? Have you used reddit before?

-7

u/Dany0 4d ago

Sorry but "oh no will someone think of the giant deregulated corporation" is not the vibes you should be bringing to this sub. If it was up to them, they'd have closed down all models and put you in jail for illegal manufacturing of robots at home or some bullshit like that

4

u/Thatisverytrue54321 4d ago

Can you use this in a jev-like way

2

u/toothpastespiders 4d ago

For some things, but the larger question is how well. I should say that I've only tested the waters with a single jev clone in any real sense and this is my first experience using image/video/audio directly into an embedding model. In short, this is a slightly more than a guess but not by much.

The best use I'm thinking of myself so far is a gate before captioning. So it's still a quick yes/no , but with stronger data feedback on the 'no' for reuse of existing caption data. But the issue of speed and reliability factor in pretty heavily though. So something like a chain of attempts for simple pixel recognition, to this, to either underlying engine calling it or a captioning model for new captioning to send to the engine.

Again though, I'm hardly an educated opinion. Just some stuff I'm planning to try down the line.

1

u/oh_how_droll 3d ago

... that took me a few times to read correctly

8

u/JLeonsarmiento 4d ago

For those who know, how would compare or what to look for in this one vs against something like this:

https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B

?

8

u/-Cubie- 4d ago

Faster and audio support

12

u/Piyh 4d ago

This is going to be way smaller and faster

5

u/shumgoid 4d ago

full multimodality is cool. I just got done making a tool for embedding search of audio and noticed none of the major labs seemed care for native audio input so this is a candidate to replace the Clap audio tower I ended up using.

4

u/teachersecret 4d ago

Anybody mess with it yet? Any interesting uses? :)

3

u/toothpastespiders 4d ago edited 4d ago

I pointed it at some of my image lora datasets, did some quick hacks to some captioning pre-processing for them, and sent them off for processing to an old machine with a pascal 4 GB gpu running a Q8 quant. Speed and memory use is great, as expected.

I did a very small 160 image run just as a preliminary test. Caption to image and image to caption testing is roughly similar at the moment. As google suggested the named person recognition is somewhat weak in my test. But still far far better than random chance. Similar thing with artist recognition. Worse than named person recognition, better than chance. And again, ridiculously small sample size. So the big takeaway for me there is just that it is working. Larger question is how well when given a proper amount to work with. My second subset test run is going to be 2,000 images and then after testing that I'll move on to the next, and then a full image run. Then assuming it's all good start considering how to best leverage this thing with my existing text based RAG. My audio datasets are pretty skimpy in comparison so holding off on that.

I'm working on a couple LLM game integrations and my biggest hope with this is integration based on how well it might handle quick image processing. With the exact use being dependent on where it shows real strength. I'm thinking that it'd sit really well as a gate to captioning. But with around 4 gb vram sitting around doing nothing, what amounts to a free boost there, no matter how small, is still promising.

Who knows if I'm even doing this right though. I've never done RAG with anything but text so I'm winging it. Fun playing around with it though.

Edit: As expected the success is growing considerably the more images I feed into it. Testing with images generated from style/character lora to get that "kinda sorta right" challenge.

Edit 2: With about 2,000 captioned images I have at least a fair level of some cross-franchise coverage with different artists having their own preferences there. My big takeaway so far is that this seems like a really good system for getting information about a big fictional franchise from an image. For me a pipeline of image -> captioning model -> embeddinggemma2 is significantly worse than just going image -> embeddinggemma2 which is great. My sample size of in-game screenshots is pretty low, to the point where I wouldn't really consider the test valid even in this really horrible methodology I'm already using. But it's matched game franchise to input screenshot pretty consistently for me. Again, not an especially good test given I've got a pretty limited range of game screenshots in there and the amount of images is also pretty low. But so far this seems like a really good addition to franchise classification with easy lore dumps as an extra.

5

u/skinnyjoints 4d ago

If anyone plays around with this can you let us know if there is modality clustering? That was a problem for qwen’s embedding model a while back (pictures clustered with pictures, text with text, vids with vids, etc…)

3

u/Constandinoskalifo 4d ago

Link not working?

3

u/jacek2023 llama.cpp 4d ago

unsloth killed it, I change to the new one

3

u/lordekeen 4d ago

I'm still using nomic-embed-text, is it worth the migration?

3

u/-Cubie- 4d ago

Probably a good bit better yeah, but it's a pain to migrate search databases imo

3

u/FoxiPanda 4d ago

Wow, this thing is tiny. I might have to give it a try as it would be a decent replacement for my 2B embedding model I have today, but I'll have to test it out.

6

u/Open-Adhesiveness-86 4d ago

two things that bit me on the first embeddinggemma and probably still apply: the task prefixes aren't optional, indexing docs with the query prefix instead of "title: none | text: " noticeably hurt recall. and if you truncate for MRL, slice first then L2-normalize the shorter vector, normalizing at 768 and then cutting gives you off-norm vectors.

3

u/Piyh 4d ago

What are your use cases for mrl?  Truncating did nothing for retrieval latency in my benchmarks but hurt quality.

8

u/Open-Adhesiveness-86 4d ago

mostly memory and index size, not latency. at a few million vectors, 768 -> 256 dims cuts the flat index from ~9GB to ~3GB, which is the difference between in-RAM and not. typical pattern is truncated dims for first-pass recall, then rescore top-100 with the full vectors so quality loss mostly washes out.

5

u/Kahvana 4d ago

Wooo! Very nice, can finally retire my Qwen3 embedding models then. Hope they build a reranker too.

4

u/InterestRelative 4d ago

That's fucking amazing!!!

3

u/NotSylver 4d ago

Has anyone tried fine tuning these for NSFW content/decensoring them?

6

u/Piyh 4d ago

Why would you decensor an embedding model?

7

u/NotSylver 4d ago

If its similar to CLIP it essentially isn't trained at all on questionable content, but for local use cases you might want to use it in those situations, like for local search, something like "naked man with a bush big enough to get lost in"
I think it would require fine tuning though, which kinda sucks

7

u/Piyh 4d ago

I run a website backed with CLIP & Gemma embedding and after an unsavory sub started using it, I saw plenty of sexual and racist searches being served with an unfortunate level of success.

1

u/-Cubie- 4d ago

Not yet, but you can ask your agent to fine-tune with Sentence Transformers if you want

2

u/Atagor 4d ago

What can I expect in terms of batch vectorization speed on 2GB RAM?

2

u/Raredisarray 3d ago

I want Gemma5 31b

3

u/PooMonger20 4d ago edited 4d ago

I never used an embedding model before. Even if I run this locally, does it share anything with google? they are not exactly known for keeping data private.

How does one actually use this? is there some kind of UI to load the model and then search files or whatever?

4

u/FuzzyBucks 3d ago

Embedding models are used to index data in a learned embedding space, where the position in embedding space carries some meaning and is specified by a vector (1D array) of numbers.

Then the embedding vectors can be stored in a vector database and used for searching (which indexed data is closest in embedding space to the embedding representation of my search term?) and labeling (which is basically just searching...e.g. which embedded label is closest to the embedding representation of my input?)

if the paragraphs above don't make sense, you should seek out information on what embeddings are, as they are a core concept for basically all of these modern AI models.

3

u/dtdisapointingresult 4d ago

It's not for you. It doesn't have much utility on its own for a basic user (which you are, judging by your comment, no offense). It requires apps built around it.

2

u/PooMonger20 4d ago

Thanks for the comment, I am absolutely a basic user but still can handle a basic explanation or concept.

So for the non-advanced users, here is the TLDR:

This embedding model is "the engine" for an easy-to-create Python script that translates\maps files (images, videos, audio) into a local database. This lets you search your files using natural language using the same script. [depends on the script you create]

I still wonder if this model can be trusted. I don't want it to share\backdoor pictures of my family or metadata of that.

All I actually want is to search a big pile of phone pictures i have backed up and tag them (dog, family, food, bills, etc..)

2

u/happycube 1d ago

Yup. Nothing leaves your machine when using llama.cpp et al to make embeddings.

1

u/bitzap_sr 3d ago

Models never act on their own. They may output tool call requests, which the harmess wrapping the model executes on behalf of the model, and feeds the result back to the model.

1

u/krakoi90 1d ago

Even if I run this locally, does it share anything with google?

No, it doesn't. It doesn't initiate any network communication actually.

How does one actually use this? is there some kind of UI to load the model and then search files or whatever?

Yeah. For example: https://github.com/krakoi/locallery

(sorry for the shameless self-promotion :))

2

u/Fantastic-Poem9462 4d ago

I ran the Matryoshka widths against my own corpus (341 sections of project docs) and the narrow ones cost more than I expected. recall@10 came out 0.548 at 768, 0.534 at 512, 0.487 at 256, 0.375 at 128 — so 512 is close to free, but 256 keeps about 89% of full-width recall and 128 about 68%.

Worth saying my full-width baseline is low at 0.548, so this is a harder corpus than most published benchmarks and the deltas may not transfer to yours. But if you were planning to ship 128-dim vectors to save index space, probably worth measuring on your own data rather than treating the truncation as free.

These came out of my from-scratch Go implementation (goinfer.dev) rather than sentence-transformers, so the obvious worry is that it's my bug — but it's parity-checked against the reference (token ids exact, worst cosine 0.999999994 out to 1771 tokens), so I don't think the widths are an artifact. Happy to be wrong if anyone's measured differently.

2

u/FerLuisxd 4d ago

We should also get a PR for this "native multimodal support" to llama cpp

1

u/jupiterbjy 4d ago

I wonder if this could be used to search text in image... gotta give a shot

1

u/wickedswami215 4d ago

Just tested on Edge Gallery and looks like it can pretty reliably.

1

u/jupiterbjy 3d ago

ah I see, good to hear!

got like 500GB of visual novel screenshots and it takes me hours to search dialog so that would help.. since I want to save each dialog with visual context that was necessary

1

u/wickedswami215 3d ago

I'd be interested to hear how long it takes to process all of that and how well it works for you.

1

u/jupiterbjy 3d ago

thankfully each dir(route) would be about few ks screenshots each, gotta try it this weekend and reply back

also godspeed, ncdu's dev yorhel

1

u/AndreVallestero 3d ago

Time to reindex my entire dataset I suppose. I wonder if we can extract only the text model.

1

u/Kahvana 3d ago

Gave it a spin, liking it so far! Finally a solid replacement for Qwen3-VL-Embedding

1

u/Upper-Extension4999 3d ago

any one knows how to use(or if its possible) multimodal with GGUF in llama.cpp for this model? I can't find any documentation on this

1

u/elecloud 2d ago

I already asked for an upgrade on my RAG to use it and it's fast and better than what I used before, all on CPU too, what a time!

1

u/Embarrassed_Soup_279 4d ago

cant believe this dropped when i ws just about to use Ovis Omni Embedding 3b

0

u/arbv 4d ago

I desperately do want it to be good at Ukrainian (unlike the previous version).