EmbeddingGemma - Micro Embeddings for Mobile Devices
Key Takeaways
The video discusses EmbeddingGemma, a 300M model for micro embeddings on mobile devices, and its applications in RAG search and semantic similarity tasks. It also covers the use of Gemma models, including Gemma 3N, T5 Gemma, and Gemma 3270M, for on-device and research applications.
Full Transcript
Okay, so over the past few months, the Gemma team has been on a roll. It's only about 6 months since they introduced Gemma 3, which was a whole suite of different models. But since then, we've seen the Gemma team actually go in two directions, which are really interesting. So the first one is that we've seen them go into making ondevice versions of the models. So at IO this year, Gemma 3N was announced around sort of June. It came out and this model is pretty amazing with what you can do on device with this. And Google's put up a whole bunch of things around developer guides for how to use this both on mobile phones, but people are also using these on Raspberry Pies and even in the browser. The other direction that we've seen the team go in is making models that can be used for various kinds of research. So, one of the ones that I didn't cover, but I found very interesting was the whole T5 Gemma. So, the T5 models go back to 2019, and there were some of the best transformer models that had both encoder and decoder. So, the encoder obviously is similar to the BERT model, the decoder being similar to the GPT models, but it was really cool that the Gemma team released a variant of these that people could use for testing out a whole bunch of different research ideas, etc. Following that, the Gemma team released Gemma 3270M. Now, this model is really fascinating because it's just so small and yet they trained this on 6 trillion tokens. So, we get to see actually what a small model can do when it's trained for a long time. Also, this model is ideal if you want to fine-tune something that's really small for a phone or for something else at the edge. And that brings us to today where the Gemma team is actually rolling out embedding Gemma. And the idea behind this is to have a really small model that can be used for doing embeddings on devices. So if you want to build like a micro rag system for a phone, for a Raspberry Pi or something, then this is going to be the model that you can use for this kind of thing. So this is actually built from uh the T5 Gemma initialization and they're releasing a variety of different versions that you can use for this that are both quantized versions I think some onyx versions some QAT versions etc. And the idea here is that you can put in text. So they're not multimodal. They're actually just text only but you can put in text up to sort of 2,00 tokens. And because they're actually matrioska embeddings you can get different dimensions of sizes out. So that can be 768 at the biggest size right down to 128 for the smallest size. So the idea with these is you've got some beddings that are really small. They can be used for a whole bunch of different applications. And here we can see that they made a hyping face space to basically show off this where you can just type in a mood or a scene or in this case as you can see like I'm putting in a product and then it will actually generate the colors that basically match this. So you can see this is what it's come out with in here. I said a sea storm. You can see sure enough we get different colors coming out of this. Okay. So the idea with these embeddings is that they're made for ondevice usage. These are really tiny, right? We're just looking at 300m in here. So to put that in perspective, the Quen embeddings that I really liked recently were almost double the size in here. And we can see here looking at their MTEBB scores just how embedding Gemma is actually comparing not only to that Quen model which is no surprise that it's beating embedding Gemma it's double the size but even when we compare to other models that are similar size or even some of the other models that are around double the size. We can see that embedding Gemma is performing really well for a whole bunch of these different tasks. whether that's retrieval, whether that's classification, or even things like clustering. So the cool thing with this is it allows you to basically do things like rag, things like semantic search, etc., all on a mobile device without you having to have access to the internet. Let's jump in and have a little play with these in code so we can see what it can actually do. All righty. So the way that you can use these embeddings if you want to use them in Python is to use the sentence transformers library. So this is similar to the transformers library. It's just made for embeddings. So you can see here that I can set it up to basically I've got a query and I've got some documents. Right? So in this case my query is which planet is known as the red planet. I've got four different documents in there and obviously our answer we'd like the answer to be Mars coming out of this. But we can see that there's some other things in there that kind of hint at a red planet as well. So we've got two that mention this. So you can see that when we run this through, we basically encode our query. We then encode the actual documents. So this is the list of strings in here. If we look at those, we can see that sure enough, our query is going to be a vector of 768 in this case. And our documents, we've got four of them, and they're also vectors of 768. So now if we come down and we basically just do model.similarity which is going to compare those vectors. So we've got the query embeddings we've got the document embeddings and it's going to print out the similarities. That's what we get back. So we can see that out of these the second one is the highest similarity in here. And then what I basically do is I can just take the documents and do an a numpy argax on this vector that we got out here. to get the position of the highest embedding and use that as the position of the string in documents that we want to print out. So we can see sure enough it comes back with Mars known for its reddish appearance is often referred to as the red planet. So that's just a simple example of doing straight up sort of comparison and stuff like that. Now you can also use this for making lots of embeddings for clustering, making embeddings for doing sort of classification etc. similar to this. Another thing if you want to try out is building different sort of little rag systems and stuff. So I put together a very simple rag in here that we're actually going to use just the Gemma models and just the sort of small ones. So in this case I'm using the Gemma 3N model. You can play around with the Gemma 3270. Both of these I was having some issues just working out, okay, what's the string supposed to be going in. This seems to work better for me at the moment. So I've just gone for this one, but you could play around with this small one as well in here. Once you load that up, you can just use that model. So obviously here, if I ask it, what is the meaning of life? We're going to get this is the ultimate question. We get a sort of response out, right? Nothing unusual there. Then we want to set up the embeddings for this. So I'm going to use langraph and lang chain in here. I've done this from the base embeddings. One of the reasons is it seems that the hugging face embeddings seems to be for an old version of sentence transformers. There's some changes that seem to be for encode document and encode query which seem to make this work. Whereas if I just use I seem to be running into issues there. From there it's very simple. We can basically set up our Chroma database. We pass in those embeddings. This is where I'm going to save the database. I then basically bring in some data. So, I've just brought in the blog post from the Gemma 3270 in here. We're loading that up, splitting it up. Nothing fancy there, just really simple sort of chunk splitting. And then I'm actually splitting it out into text. I add the text to the vector store. You can see that if I do some searches, it's looking like, okay, it's starting to find some of the right information. And then I basically just want to build a simple LAN graph for this. So we're just going to have if we look at the sort of print out we're just going to have a start a retrieve section where we'll go to the vector store get the information using the embedding Gemma model in there and then come back and generate out using the Gemma 3N model here. And you can see that here that we've got our retrieve node which is basically just going to the vector store getting the question that we've passed in retrieving that. We've then got the generate where we're constructing a prompt to pass it in. I'm just passing it directly into this. So I'm not actually using the lang chain llm kind of model system here. I'm just basically using the hanging face one passing it in getting it back. Saving that in. And then we basically just add these together and compile it. So here you can see if we ask it basically what is Gemma 270? Sure enough it gives us a response back. JA 3270 is a compact model for hyperefficient AI. It then gives us some other information. If we come down and look at the response, we can see these are not great LLMs and I probably should look into setting up the ML chat format for them to get better results out. But you can see that it certainly got what we wanted. So what is Gemma 270M? Based on the context, Gemma 270M is a compact model for high efficient AI. It's part of the Gemma series of models designed for text generation, chat bots, code generation, creative writing, sentiment analysis. It gives us a whole bunch of list of things in there. It's suitable for ondevice and research applications due to its size and performance. So, it's certainly been able to do that and it's used the embeddings in there to do this. Now, are these the necessarily the best embeddings that you would use on a server, etc.? No, you can see here though just how much RAM that I'm actually using for those models. It's just tiny amounts of VRAM on the GPU in here. The embeddings themselves actually run quite quickly just off the CPU. So you can get some good results just running like this first bit up here. No problems just using the CPU in this. So clearly these are made for ondevice use. And if you are looking at doing something like that, you can definitely set these up to run on a mobile phone and create embeddings for a whole bunch of different things. Earlier on, I showed you the mood pallet one. If we come in and look at the code for that, we can actually see that really what this is doing is just doing an embedding of whatever you put in versus these descriptions and then returning the name and the hex format back from this. So you can see there's lots of interesting things that you could use this on a phone for games for a variety of different things. And because it's so small, it pretty much will run on most phones that are out there to date. It doesn't need to be the latest, most fancy phone. And when you combine this with the other Gemma 3 N series of models, etc., then you've got both the LLM and the embeddings to actually make some kind of rag system or something. Okay. So if you're doing any work with mobile or on the edge sort of stuff, definitely check out these models. It is a really interesting direction how the Gemma series is splitting into sort of one path being on device stuff at the moment and one path being research focused kind of things. It will be interesting to see when we get a new sort of fuller version of Gemma how they go with sizes and stuff like that. Anyway, as always, if you found the video useful, please click like and subscribe, and I will talk to you in the next video. Bye for now.
Original Description
In this video, I look at the latest Gemma release which is EmbeddingGemma, a 300M model that can be used for doing embedding tasks like RAG and semantic similarity on phones and mobile devices at the edge.
Blog: https://developers.googleblog.com/en/introducing-embeddinggemma/
colab: https://dripl.ink/nTzXV
For more tutorials on using LLMs and building agents, check out my Patreon
Patreon: https://www.patreon.com/SamWitteveen
Twitter: https://x.com/Sam_Witteveen
🕵️ Interested in building LLM Agents? Fill out the form below
Building LLM Agents Form: https://drp.li/dIMes
👨💻Github:
https://github.com/samwit/llm-tutorials
⏱️Time Stamps:
00:00 Intro
00:09 Gemma 3
00:36 Gemma 3n
00:47 T5Gemma
01:24 Gemma 3 270M
01:51 EmbeddingGemma
04:21 Colab Demo
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Sam Witteveen · Sam Witteveen · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
LangChain Basics Tutorial #1 - LLMs & PromptTemplates with Colab
Sam Witteveen
LangChain Basics Tutorial #2 Tools and Chains
Sam Witteveen
ChatGPT API Announcement & Code Walkthrough with LangChain
Sam Witteveen
Trying Out Flan 20B with UL2 - Working in Colab with 8Bit Inference
Sam Witteveen
LangChain - Conversations with Memory (explanation & code walkthrough)
Sam Witteveen
LangChain Chat with Flan20B
Sam Witteveen
LangChain - Using Hugging Face Models locally (code walkthrough)
Sam Witteveen
PAL : Program-aided Language Models with LangChain code
Sam Witteveen
Building a Summarization System with LangChain and GPT-3 - Part 1
Sam Witteveen
Building a Summarization System with LangChain and GPT-3 - Part 2
Sam Witteveen
Microsoft's Visual ChatGPT using LangChain
Sam Witteveen
Building a Summarization System with LangChain - Part 3 Using ChatGPT Turbo
Sam Witteveen
LangChain Agents - Joining Tools and Chains with Decisions
Sam Witteveen
Investigating Alpaca 7B - Finetuned LLaMa LLM
Sam Witteveen
Comparing LLMs with LangChain
Sam Witteveen
Running Alpaca7B in Colab
Sam Witteveen
How to finetune your own Alpaca 7B
Sam Witteveen
How to make a custom dataset like Alpaca7B
Sam Witteveen
Understanding Constitutional AI - the paper and key concepts
Sam Witteveen
Using Constitutional AI in LangChain
Sam Witteveen
Talking to Alpaca with LangChain - Creating an Alpaca Chatbot
Sam Witteveen
Text-to-video-synthesis with Diffusers and Colab
Sam Witteveen
Meet Dolly the new Alpaca model
Sam Witteveen
Checking out the Cerebras-GPT family of models
Sam Witteveen
A Step-by-Step Guide to Fine-Tuning Your Dolly Model (tutorial)
Sam Witteveen
Is GPT4All your new personal ChatGPT?
Sam Witteveen
Raven - RWKV-7B RNN's LLM Strikes Back
Sam Witteveen
Talk to your CSV & Excel with LangChain
Sam Witteveen
Vicuna - 90% of ChatGPT quality by using a new dataset?
Sam Witteveen
Koala Revealed: The ChatGPT Alternative You Need to Know! 🔍
Sam Witteveen
Running Koala for free in Colab. Your own personal ChatGPT? (tutorial)
Sam Witteveen
BabyAGI: Discover the Power of Task-Driven Autonomous Agents!
Sam Witteveen
Auto-GPT - How to Automate a Task Based AI with GPT-4
Sam Witteveen
Improve your BabyAGI with LangChain
Sam Witteveen
Generative Agents - Deep Dive and GPT-4 Recreation
Sam Witteveen
GPT4ALLv2: The Improvements and Drawbacks You Need to Know!
Sam Witteveen
Dolly 2.0 by Databricks: Open for Business but is it Ready to Impress!
Sam Witteveen
Red Pajama - Operation: Freeing LLaMA
Sam Witteveen
Investigating Open Assistant - Models, Datasets and Addons
Sam Witteveen
Investigating MiniGPT-4 - The Secret behind GPT-V?
Sam Witteveen
Stable LM 3B - The new tiny kid on the block.
Sam Witteveen
Bard can now code and put that code in Colab for you.
Sam Witteveen
Checking out Bark: a Text to Speech system by Suno AI
Sam Witteveen
Fine-tuning LLMs with PEFT and LoRA
Sam Witteveen
Master PDF Chat with LangChain - Your essential guide to queries on documents
Sam Witteveen
Using LangChain with DuckDuckGO Wikipedia & PythonREPL Tools
Sam Witteveen
Building Custom Tools and Agents with LangChain (gpt-3.5-turbo)
Sam Witteveen
StableVicuna: The New King of Open ChatGPTs?
Sam Witteveen
WizardLM: Evolving Instruction Datasets to Create a Better Model
Sam Witteveen
LaMini-LM - Mini Models Maxi Data!
Sam Witteveen
Finding the Best Free ChatGPT
Sam Witteveen
MPT-7B - The First Commercially Usable Fully Trained LLaMA Style Model
Sam Witteveen
LangChain Retrieval QA Over Multiple Files with ChromaDB
Sam Witteveen
LangChain Retrieval QA with Instructor Embeddings & ChromaDB for PDFs
Sam Witteveen
LangChain + Retrieval Local LLMs for Retrieval QA - No OpenAI!!!
Sam Witteveen
Transformers Agent - Is this Hugging Face's LangChain Competitor?
Sam Witteveen
StarCoder - The LLM to make you a coding star?
Sam Witteveen
Testing Starcoder for Reasoning with PAL
Sam Witteveen
The New Wizards - Unfiltered & Unaligned
Sam Witteveen
Camel + LangChain for Synthetic Data & Market Research
Sam Witteveen
More on: RAG Basics
View skill →Related Reads
📰
📰
📰
📰
What Is RAG AI? Retrieval-Augmented Generation Explained — American Dream AI
Medium · AI
Treat Retrieved Content as Data, Not Instructions
Dev.to AI
How to Evaluate Production RAG: Keyword, Vector, SQL, and Hybrid Retrieval
Dev.to · Anya Summers
RAG Database Design: SQL, Full-Text Search, Vector Search, and Context Retrieval
Dev.to · puffball1567
Chapters (7)
Intro
0:09
Gemma 3
0:36
Gemma 3n
0:47
T5Gemma
1:24
Gemma 3 270M
1:51
EmbeddingGemma
4:21
Colab Demo
🎓
Tutor Explanation
DeepCamp AI