Deploy Hugging Face models from Vertex AI Model Garden
Key Takeaways
Deploying Hugging Face models from Vertex AI Model Garden using Google Cloud Web console and Kubernetes Engine, leveraging Hugging Face Hub for model access and management.
Full Transcript
[Music] you can deploy thousands of AI models with Google cloud and hugging face without leaving the Google Cloud Web console and without setting up complex infrastructure I'm sure you want to know how so let's get started this is the Google Cloud Web console let me use the search bar to go to vertx AI now in vertex ai go to the model garden and has many machine learning model the models that are available in the model Garden can handle various types of data including language images and table data they can also perform a wide range of tasks including generating text classifying data and translating languages so you can find models from Google here and from other providers and one of those providers is hugging face let's explore here's a list of thousands of machine learning models straight from the face up you might see some models here that you're familiar with here's llama there's mistr and there's Google's Gemma too so let's deploy Gemma you'll need to fill in a hooking face access token now let me show you how to create a token you'll need to do this only once so first go to the hugging face Hub hugging face. and create an account once you're logged in choose settings then access tokens and click create a new token give the token a name and allow a read access to all public gated repositories you can access make sure to copy the token and you also need to accept the Gemma license right search for Gemma 2 and acknowledge the license agreement now back to vertex AI fill in your hugging face access token click deploy wait for a while and make a prediction request alternative ly you can also deploy using Google kubernetes engine make sure you have a gpe autopilot cluster with gpus ready or if you're using standard enable note out of provisioning then copy and apply the following manifest it creates a deployment with the hugging face deep learning containers and a service to expose the deployment it also creates a secret with your hugging phase token because it needs to pull those models from up so that's that you can thousands of models from the hugging face up without leaving the Google Cloud Web console [Music]
Original Description
Model Garden → https://goo.gle/3zXNQ5P
Explore AI models in Model Garden → https://goo.gle/3Yn4zJ2
Hugging Face models → https://goo.gle/400RczH
You can deploy thousands of AI models from the Hugging Face Hub directly through the Google Cloud web console without setting up complex infrastructure. Vertex AI's Model Garden provides access to many models, including those from Hugging Face, supporting diverse data types and tasks. Watch along as Googler, Wietse Venema, walks through deploying a HuggingFace open model on Google Cloud’s Model Garden.
Chapters:
0:00 - Intro
0:19 - Get started with Vertex AI Model Garden
0:55 - Deploying HuggingFace models on Vertex AI
2:27 - Get started today
Watch more Google Cloud: Building with Hugging Face → https://goo.gle/BuildWithHuggingFace
Subscribe to Google Cloud Tech → https://goo.gle/GoogleCloudTech
#HuggingFace #GoogleCloud
Speaker: Wietse Venema
Products Mentioned: Vertex AI, Model Garden, Open Source, Gemini
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Google Cloud Tech · Google Cloud Tech · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
I’m going for it #GoogleCloudCertified
Google Cloud Tech
I had to get #GoogleCloudCertified
Google Cloud Tech
Be better overall at what you do #GoogleCloudCertified
Google Cloud Tech
Cloud Monitoring on our radar #Analysis #Uptime
Google Cloud Tech
Introduction to Generative AI Studio
Google Cloud Tech
How to use Github Actions with Google's Workload Identity Federation
Google Cloud Tech
Introduction to Responsible AI
Google Cloud Tech
Networking updates and CDMC-certified architecture
Google Cloud Tech
Create and use a Cloud Storage bucket
Google Cloud Tech
How to digitize text from documents
Google Cloud Tech
Faster analytical queries with AlloyDB
Google Cloud Tech
Next ‘23 sessions and FaaS Wave
Google Cloud Tech
Introduction to Assured Open Source Software
Google Cloud Tech
BigQuery Cost Optimization: Storage
Google Cloud Tech
BigQuery Cost Optimization: Compute
Google Cloud Tech
BigQuery Cost Optimization: Select Queries
Google Cloud Tech
Remote Field Equipment Management with Manufacturing Data Engine
Google Cloud Tech
Supercharging your applications with Cloud SQL Enterprise Plus
Google Cloud Tech
Vector Support on our radar #GenAI
Google Cloud Tech
Architecting a blockchain startup with Google Cloud
Google Cloud Tech
Kubernetes and multitasking updates!
Google Cloud Tech
GKE: Using Kubernetes Events
Google Cloud Tech
How to configure firewall rules for Cloud Composer
Google Cloud Tech
Vertex AI Embeddings API + Matching Engine: Grounding LLMs made easy
Google Cloud Tech
Geospatial analytics on our radar #EarthEngine #BigQuery
Google Cloud Tech
Ensuring requests are set in Kubernetes
Google Cloud Tech
Cloud Next 2023, Google research program, and more!
Google Cloud Tech
How to migrate projects between organizations with Resource Manager
Google Cloud Tech
How to run #MySQL in Google Cloud
Google Cloud Tech
#GenerativeAI for enterprises and #Next2023
Google Cloud Tech
How Google Photos scales to store 4 trillion photos and videos
Google Cloud Tech
Google Cross-Cloud Interconnect (Demo 2)
Google Cloud Tech
GKE Cost Optimization Golden Signals: Introduction
Google Cloud Tech
GKE Cost Optimization Golden Signals: Workload Rightsizing
Google Cloud Tech
GKE Load Balancing: Overview
Google Cloud Tech
GKE Load Balancing: Best Practices
Google Cloud Tech
Disaster Recovery in GKE
Google Cloud Tech
How to configure IP masquerade agent in GKE Standard clusters
Google Cloud Tech
Enable and use GKE Control plane logs
Google Cloud Tech
Compliance in Australia with Assured Workloads
Google Cloud Tech
Creating budgets and budget alerts in Google Cloud #FinOps
Google Cloud Tech
Cloud SQL Enterprise Plus on our radar #mySQL
Google Cloud Tech
What's Next for Google Cloud?
Google Cloud Tech
How Loveholidays scaled with Contact Center AI
Google Cloud Tech
What is fleet team management in GKE?
Google Cloud Tech
Troubleshoot VPC Network Peering
Google Cloud Tech
Introduction to DocAI and Contact Center AI
Google Cloud Tech
Cloud Run Direct VPC egress explained
Google Cloud Tech
Database deployment options in GKE
Google Cloud Tech
Analyze cloud billing data with #BigQuery
Google Cloud Tech
Tips to becoming a world-class Prompt Engineer
Google Cloud Tech
Serverless is simple. Do I need CI/CD?
Google Cloud Tech
Accelerating model deployment with MLOps
Google Cloud Tech
How Hawaii's Department of Human Services scaled with CCAI
Google Cloud Tech
Pricing API on our #Radar
Google Cloud Tech
How Recommendations AI for Media can boost customer retention
Google Cloud Tech
Troubleshooting: Node Not Ready Status
Google Cloud Tech
One weekend until Cloud Next 2023!
Google Cloud Tech
#GoogleCloudNext starts tomorrow!
Google Cloud Tech
#GoogleCloudNext will be demand!
Google Cloud Tech
More on: LLMOps
View skill →Related Reads
📰
📰
📰
📰
My LLM drift tracker flagged four regressions this week. All four were wrong.
Dev.to AI
I built a memory layer that works across Claude, ChatGPT and Cursor — and rendered it as a 3D brain
Dev.to AI
Put an LLM QA gate in front of your paid deliverable: fail-closed review, one self-heal, no retry loops
Dev.to · Kynth
Jailbreaking de LLMs mediante Prompt Injection
Medium · AI
Chapters (4)
Intro
0:19
Get started with Vertex AI Model Garden
0:55
Deploying HuggingFace models on Vertex AI
2:27
Get started today
🎓
Tutor Explanation
DeepCamp AI