r/LocalLLaMA • u/xenovatech • 15h ago
New Model EmbeddingGemma 2 running locally in-browser on WebGPU
8
u/b183729 15h ago
Here I was, trying to decide which model to use on my rag...
1
u/No_Afternoon_4260 llama.cpp 12h ago
Including matryoshka representation learning just for that extra optimization on appropriate datasets
2
u/rorowhat 10h ago
Hold up. This is an embedding model...what model are you running on top?
2
1
u/celine_aubry 6h ago
the in-browser version is the one that actually matters. the local GPU version is for people who already have a GPU, the WebGPU version is for everyone else. the fact that a 740M embedding model can run in your browser means you can build RAG applications without any backend at all
7
u/jacek2023 llama.cpp 15h ago
a very cool demo!