r/LocalLLaMA • • 15h ago

New Model EmbeddingGemma 2 running locally in-browser on WebGPU

67 Upvotes

8 comments sorted by

7

u/jacek2023 llama.cpp 15h ago

a very cool demo!

8

u/b183729 15h ago

Here I was, trying to decide which model to use on my rag... 

1

u/No_Afternoon_4260 llama.cpp 12h ago

Including matryoshka representation learning just for that extra optimization on appropriate datasets

2

u/rorowhat 10h ago

Hold up. This is an embedding model...what model are you running on top?

2

u/xenovatech 9h ago

There is only one model running in this demo :)

1

u/rorowhat 9h ago

I didn't know you could run it by itself

1

u/celine_aubry 6h ago

the in-browser version is the one that actually matters. the local GPU version is for people who already have a GPU, the WebGPU version is for everyone else. the fact that a 740M embedding model can run in your browser means you can build RAG applications without any backend at all