Deploy Gemma 2 with multiple LoRA adapters on GKE
Key Takeaways
This video demonstrates how to deploy Gemma 2 with multiple LoRA adapters using TGI on Google Kubernetes Engine (GKE), allowing for dynamic selection of adapters for incoming requests.
Full Transcript
[Music] hi I'm V from Google cloud and I'm abdell today we'll show you how to deploy Gemma 2 with multiple Dora adapters using TGI on Google kubernetes engine that sounds complicated can you break it down for us sure let's go over some terms first Gema is Google's open large language model and TGI is an open source Next Generation inference server from huging phase now low rank adaptation or Laura is a fine-tuning technique used to adapt a base model to specific tasks without retraining the entire model instead Laura introduces small adapter layers that are trained on a targeted data set while the original base model remains Frozen so why would you want to deploy Gemma 2 with multiple Laura adapters on jaki well there are a few benefits in a multi Laura deployment you deploy this single base model and then load multiple Laura adapters and you can dynamically select a different adapter for each incoming requests so you don't have to create multiple deployments which is more cost-efficient that sounds great how do you actually deploy Gemma 2 with multiple Laura adapters with TGI DLC on GK um I was hoping that you would show me that sure well it's a multi-step process first you need to create a GK cluster we already covered this in previous videos check the links in the description box after that you need to deploy Gemma 2 with multiple adapters using the TGI deep learning container on the cluster that you're right that doesn't sound complicated at all is there anything else people should know well first of all let's look at some y files first we have my deployment. yl file this file describes our deployment including the number of replicas the container images we are using in this case TGI and any environment variables see this section here this is where we specify the GMA 2 model and theora adapter we want to deploy next we have our Ingress yaml these files Define how our deployed container will be exposed to the outside world it basically creates a load balancer so we can access our Gemma 2 model enough enough jaml files let's get this show on the road can you show me how to actually send a request to it sure using curl you can send request to the model and specify the adapter layer to use in the particular case I am using the SQL adapter okay it looks like you got a response back but I was hoping you can show me how a developer would use this in the r of course using the hugging face Hub is dek I have to import the inference client specify the model endpoint that is the IP address of my load balancer and invoke the model using the magic coder adapter layer to generate some Rust code that's amazing this is this is rust code so you're saying that this is one deployment of a base model and then TGI selects a different adapter based on the request that's pretty much it awesome well there you have it we showed you how to deploy a Gemma 2 based model with multiple Laura adapters don't forget to check out the description for links to all the resources we use today including those yaml files and if you have any question drop them in the comments below thanks for watching and we'll see you in the next one [Music]
Original Description
Tutorial: Deploy Gemma 2 with multiple LoRA adapters using TGI on GKE → https://goo.gle/4f5KP1C
Video: Train a LoRA adapter with your own dataset → https://goo.gle/4gkBLar
Deep dive: A conceptual overview of Low-Rank Adaptation (LoRA) → https://goo.gle/4in4NrA
Learn how to deploy multiple LoRA adapters in one deployment on Google Kubernetes Engine. Low-Rank Adaptation, or LoRA, is a fine-tuning technique used to adapt a base model to specific tasks without retraining the entire model. Watch along and learn how to deploy Gemma 2, a powerful open large language model, and TGI, an open-source LLM inference server from Hugging Face, to deploy multiple LoRA adapters for different tasks.
More resources:
Docs: Hugging Face Hub Inference client → https://goo.gle/3Zrwo2c
Docs: An overview of the TGI command line interface flags → https://goo.gle/41Fs1nd
Watch more Google Cloud: Building with Hugging Face → https://goo.gle/BuildWithHuggingFace
Subscribe to Google Cloud Tech → https://goo.gle/GoogleCloudTech
Speakers: Wietse Venema, Abdel Sghiouar
Products Mentioned: Gemma, Gemini, Google Kubernetes Engine (GKE)
#GoogleCloud #HuggingFace
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Google Cloud Tech · Google Cloud Tech · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
I’m going for it #GoogleCloudCertified
Google Cloud Tech
I had to get #GoogleCloudCertified
Google Cloud Tech
Be better overall at what you do #GoogleCloudCertified
Google Cloud Tech
Cloud Monitoring on our radar #Analysis #Uptime
Google Cloud Tech
Introduction to Generative AI Studio
Google Cloud Tech
How to use Github Actions with Google's Workload Identity Federation
Google Cloud Tech
Introduction to Responsible AI
Google Cloud Tech
Networking updates and CDMC-certified architecture
Google Cloud Tech
Create and use a Cloud Storage bucket
Google Cloud Tech
How to digitize text from documents
Google Cloud Tech
Faster analytical queries with AlloyDB
Google Cloud Tech
Next ‘23 sessions and FaaS Wave
Google Cloud Tech
Introduction to Assured Open Source Software
Google Cloud Tech
BigQuery Cost Optimization: Storage
Google Cloud Tech
BigQuery Cost Optimization: Compute
Google Cloud Tech
BigQuery Cost Optimization: Select Queries
Google Cloud Tech
Remote Field Equipment Management with Manufacturing Data Engine
Google Cloud Tech
Supercharging your applications with Cloud SQL Enterprise Plus
Google Cloud Tech
Vector Support on our radar #GenAI
Google Cloud Tech
Architecting a blockchain startup with Google Cloud
Google Cloud Tech
Kubernetes and multitasking updates!
Google Cloud Tech
GKE: Using Kubernetes Events
Google Cloud Tech
How to configure firewall rules for Cloud Composer
Google Cloud Tech
Vertex AI Embeddings API + Matching Engine: Grounding LLMs made easy
Google Cloud Tech
Geospatial analytics on our radar #EarthEngine #BigQuery
Google Cloud Tech
Ensuring requests are set in Kubernetes
Google Cloud Tech
Cloud Next 2023, Google research program, and more!
Google Cloud Tech
How to migrate projects between organizations with Resource Manager
Google Cloud Tech
How to run #MySQL in Google Cloud
Google Cloud Tech
#GenerativeAI for enterprises and #Next2023
Google Cloud Tech
How Google Photos scales to store 4 trillion photos and videos
Google Cloud Tech
Google Cross-Cloud Interconnect (Demo 2)
Google Cloud Tech
GKE Cost Optimization Golden Signals: Introduction
Google Cloud Tech
GKE Cost Optimization Golden Signals: Workload Rightsizing
Google Cloud Tech
GKE Load Balancing: Overview
Google Cloud Tech
GKE Load Balancing: Best Practices
Google Cloud Tech
Disaster Recovery in GKE
Google Cloud Tech
How to configure IP masquerade agent in GKE Standard clusters
Google Cloud Tech
Enable and use GKE Control plane logs
Google Cloud Tech
Compliance in Australia with Assured Workloads
Google Cloud Tech
Creating budgets and budget alerts in Google Cloud #FinOps
Google Cloud Tech
Cloud SQL Enterprise Plus on our radar #mySQL
Google Cloud Tech
What's Next for Google Cloud?
Google Cloud Tech
How Loveholidays scaled with Contact Center AI
Google Cloud Tech
What is fleet team management in GKE?
Google Cloud Tech
Troubleshoot VPC Network Peering
Google Cloud Tech
Introduction to DocAI and Contact Center AI
Google Cloud Tech
Cloud Run Direct VPC egress explained
Google Cloud Tech
Database deployment options in GKE
Google Cloud Tech
Analyze cloud billing data with #BigQuery
Google Cloud Tech
Tips to becoming a world-class Prompt Engineer
Google Cloud Tech
Serverless is simple. Do I need CI/CD?
Google Cloud Tech
Accelerating model deployment with MLOps
Google Cloud Tech
How Hawaii's Department of Human Services scaled with CCAI
Google Cloud Tech
Pricing API on our #Radar
Google Cloud Tech
How Recommendations AI for Media can boost customer retention
Google Cloud Tech
Troubleshooting: Node Not Ready Status
Google Cloud Tech
One weekend until Cloud Next 2023!
Google Cloud Tech
#GoogleCloudNext starts tomorrow!
Google Cloud Tech
#GoogleCloudNext will be demand!
Google Cloud Tech
More on: LLM Engineering
View skill →Related Reads
📰
📰
📰
📰
Unlocking Open-Source AI: 5 Tools for Unbeatable Privacy and Cost Efficiency
Medium · LLM
The AEO tricks that don’t work in AI search
Medium · AI
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics
Dev.to AI
Building AI Data Pipelines — How to Feed Your LLM Fresh Web Data
Dev.to AI
🎓
Tutor Explanation
DeepCamp AI