Introducing EmbeddingGemma: The Best-in-Class Open Model for On-Device Embeddings
Key Takeaways
Introduces EmbeddingGemma, a state-of-the-art open model for on-device embeddings, and discusses its capabilities and applications
Full Transcript
[Music] Hi, I'm Alice. >> I'm Lucas. We're product managers at Google DeepMind. >> And today we are incredibly excited to introduce Embedding Gemma, our state-of-the-art embedding model designed for mobile first AI. Embedding Gemma is a 300 million parameter text embedding model designed to power generative AI experiences directly on your hardware. Embeddings are numerical representations of data. This model transforms text like messages, emails or notes into a vector of numbers to represent meaning in a highdimensional space that a generative model can then use for downstream tasks. Embedding Gemma is small, fast, and efficient. Thanks to quantization aware training, you can run the model with as little as 300 megabytes of RAM while preserving state-of-the-art quality. It generates embeddings of 768 dimensions, but thanks to MROSKA representation learning, you can customize the model's output dimensions and go down to 128. Based on the same technology and research that powers our Gemini embedding models, embedding Gemma brings that state-of-the-art capability in a smaller and more lightweight model. Think highquality semantic search, fast and relevant information retrieval, or customized classification and clustering, just to name a few opportunities. Embedding Gemma achieves the best score on the comprehensive massive text embedding benchmark for models under 500 million parameters. The gold standard for text embedding evaluation trained across 100 plus languages. Embedding Gemma brings proven performance to instantly connect with diverse and global audiences. We've engineered embedding Gemma specifically for ondevice performance to ensure efficient computations and minimal memory footprint even on resource constrained hardware. Embedding Gemma facilitates ondevice embedding of local documents. So sensitive user data never leaves the device. And because it works offline, it means Frontier search and retrieval features work regardless of connectivity. Together with our generative models like Gemma 3N, you can build powerful mobile first generative AI experiences and efficient retrieval augmented generation pipelines. This means your applications can now leverage user context from data to provide more personalized and helpful responses such as understanding that you need your carpenters's number for help with damaged floorboards. Here's an example of what embedding Gemma can power. What you are seeing is how a user can utilize embedding Gemma to query previously opened articles or other web pages. The model embeds each page as it's opened in real time. Then with a browser extension that uses embedding Gemma, the user can ask a question to retrieve the contextually relevant articles. And because the embeddings are created on device, all this is happening without leaving the user's hardware. >> And it's designed with customization in mind. fine-tune embedding Gemma for your domain or in a particular language. It works across popular tools and platforms such as hugging face and Kaggle. Check out our notebook examples part of the Gemma cookbook to get started. Our next generation of ondevice embedding models is here and it's open for everyone. It's small, fast, and efficient. Download Embedding Gemma and get started building right now. >> You can find links in the description below. We can't wait to see what Embedding Gemma unlocks for you. [Music]
Original Description
Discover EmbeddingGemma, a state-of-the-art 308 million parameter text embedding model designed to power generative AI experiences directly on your hardware. Ideal for mobile-first Al, EmbeddingGemma brings powerful capabilities to your applications, enabling features like semantic search, information retrieval, and custom classification – all while running efficiently on-device.
In this video, Alice Lisak and Lucas Gonzalez from the Gemma team introduce EmbeddingGemma and explain how it works. Learn how you can run this model on less than 200MB of RAM with quantization, customize its output dimensions with Matryoshka Representation Learning (MRL), and
build powerful offline Al features.
Resources:
Learn about EmbeddingGemma → https://developers.googleblog.com/en/introducing-embeddinggemma
EmbeddingGemma documentation → https://ai.google.dev/gemma/docs/embeddinggemma
Gemma Cookbook → https://github.com/google-gemini/gemma-cookbook
Quickstart RAG notebook → https://github.com/google-gemini/gemma-cookbook/blob/main/Gemma/%5BGemma_3%5DRAG_with_EmbeddingGemma.ipynb
Discover Gemma models → https://deepmind.google/models/gemma
Chapters
0:00 - Intro
0:26 - Model overview
1:18 - Model features
2:29 - RAG
2:54 - Website embedding demo
3:23 - Tools and platforms
3:41 - Conclusion
Subscribe to Google for Developers → https://goo.gle/developers
Speaker:Alice Lisak Lucas Gonzalez
Products Mentioned: Google AI, Gemma,Generative AI
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Google for Developers · Google for Developers · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Developer Journey - Sunnyvale DSC Summit ‘19
Google for Developers
How Google is working with students - Sunnyvale DSC Summit ‘19
Google for Developers
Starting your career in the Cloud - Sunnyvale DSC Summit ‘19
Google for Developers
The Solution Challenge - Sunnyvale DSC Summit ‘19
Google for Developers
Firebase - Sunnyvale DSC Summit ‘19
Google for Developers
Cloud Hero - Sunnyvale DSC Summit ‘19
Google for Developers
Panel discussion - Sunnyvale DSC Summit ‘19
Google for Developers
The art of negotiation - Sunnyvale DSC Summit ‘19
Google for Developers
Courage to care, solve and share - Sunnyvale DSC Summit ‘19
Google for Developers
Version 9 of Angular, Glass Enterprise Edition 2, path to DX deprecation, & more!
Google for Developers
[DEPRECATING] Introducing a new series (Assistant for Developers Pro Tips)
Google for Developers
Detecting memory bugs with HWASan, Bazel 2.1, Next ‘20 session guide, & more!
Google for Developers
Why Podcast.app chose a .app domain name
Google for Developers
Machine Learning Bootcamp Jakarta 2019
Google for Developers
Android Studio 3.6, Android 11 Developer Preview, Kubeflow 1.0, & more!
Google for Developers
[DEPRECATING] Importance of community (Assistant on Air)
Google for Developers
Why the Flutter team switched from .io to a .dev domain name
Google for Developers
3 website-building tips from .dev creators
Google for Developers
Why NimbleDroid chose a .app domain name
Google for Developers
Android Platform Codelab, Bazel 2.2, Maps Android Utility Library v1.0, & more!
Google for Developers
Google for Games Developer Summit: A free, digital experience for game developers
Google for Developers
Inspecting Home Graph (Assistant for Developers Pro Tips)
Google for Developers
Google for Games Developer Summit Keynote
Google for Developers
Stadia Games & Entertainment presents: Keys to a great game pitch (Google Games Dev Summit)
Google for Developers
Empowering game developers with Stadia R&D (Google Games Dev Summit)
Google for Developers
Supercharging discoverability with Stadia (Google Games Dev Summit)
Google for Developers
Stadia Games & Entertainment presents: Creating for content creators (Google Games Dev Summit)
Google for Developers
Bringing Destiny to Stadia: A postmortem (Google Games Dev Summit)
Google for Developers
Live Captioning in Google Slides
Google for Developers
[DEPRECATING] User engagement for the Google Assistant
Google for Developers
TensorFlow Dev Summit ‘20, Google for Games Dev Summit, Cloud AI Platform Pipelines, & much more!
Google for Developers
Top 5 from the TensorFlow Dev Summit 2020
Google for Developers
Developer Student Clubs 2019 Turkey Leads Summit
Google for Developers
Building simpler payment experiences | Google Pay Plugin for Magento 2
Google for Developers
Become A Developer Student Club Lead
Google for Developers
Firebase Kotlin Extensions, ARM apps on the Android Emulator, Angular v9.1, & more!
Google for Developers
Test suite for Smart Home (Assistant for Developers Pro Tips)
Google for Developers
Google Play updates, Bazel 3.0, Business Console for Google Pay, & more!
Google for Developers
How to use error logs (Assistant for Developers Pro Tips)
Google for Developers
Contact Center AI, Android Studio 4.1 Canary 5, TensorFlow QAT API, & more!
Google for Developers
WebView DevTools, Kotlin meets gRPC, Flutter CodePen support, & more! (Episode 200)
Google for Developers
Offline handling for Smart Home (Assistant for Developers Pro Tips)
Google for Developers
Android 11 Dev Preview 3, Google Fonts for Flutter, Shielded VM, & more!
Google for Developers
Machine Learning Foundations: Ep #1 - What is ML?
Google for Developers
Flutter web support updates, BigQuery materialized views, Cloud Spanner emulator, & more!
Google for Developers
Computer vision by building a neural network with TensorFlow | Machine Learning Foundations
Google for Developers
Machine Learning Foundations: Ep #3 - Convolutions and pooling
Google for Developers
Android 11 Beta plans, Flutter 1.17, Dart 2.8, & much more!
Google for Developers
Machine Learning Foundations: Ep #4 - Coding with Convolutional Neural Networks
Google for Developers
Google Developers ML Summit
Google for Developers
Real-world image classification using convolutional neural networks | Machine Learning Foundations
Google for Developers
Adobe XD support for Flutter, Architecture Framework, temporary closures with Places API, & more!
Google for Developers
Machine Learning Foundations: Ep #6 - Convolutional cats and dogs
Google for Developers
Machine Learning Foundations: Ep #7 - Image augmentation and overfitting
Google for Developers
Announcing Firebase Live, Flutter Day, Java 11 on Google Cloud Functions, & more!
Google for Developers
Machine Learning Foundations: Ep #8 - Tokenization for Natural Language Processing
Google for Developers
Android 11 Beta, Google Play Asset Delivery, Firebase Crashlytics SDK, & much more!
Google for Developers
Natural Language Processing: Using sequencing APIs in TensorFlow | Machine Learning Foundations
Google for Developers
Build a sarcasm classifier using NLP and TensorFlow | Machine Learning Foundations
Google for Developers
AR Realism with the ARCore Depth API
Google for Developers
More on: Multimodal LLMs
View skill →Related Reads
📰
📰
📰
📰
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to AI
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to · Imus
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to AI
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to · Imus
Chapters (7)
Intro
0:26
Model overview
1:18
Model features
2:29
RAG
2:54
Website embedding demo
3:23
Tools and platforms
3:41
Conclusion
🎓
Tutor Explanation
DeepCamp AI