Deploy Hugging Face models from Vertex AI Model Garden

Google Cloud Tech · Beginner ·🧠 Large Language Models ·1y ago

Key Takeaways

Deploying Hugging Face models from Vertex AI Model Garden using Google Cloud Web console and Kubernetes Engine, leveraging Hugging Face Hub for model access and management.

Full Transcript

[Music] you can deploy thousands of AI models with Google cloud and hugging face without leaving the Google Cloud Web console and without setting up complex infrastructure I'm sure you want to know how so let's get started this is the Google Cloud Web console let me use the search bar to go to vertx AI now in vertex ai go to the model garden and has many machine learning model the models that are available in the model Garden can handle various types of data including language images and table data they can also perform a wide range of tasks including generating text classifying data and translating languages so you can find models from Google here and from other providers and one of those providers is hugging face let's explore here's a list of thousands of machine learning models straight from the face up you might see some models here that you're familiar with here's llama there's mistr and there's Google's Gemma too so let's deploy Gemma you'll need to fill in a hooking face access token now let me show you how to create a token you'll need to do this only once so first go to the hugging face Hub hugging face. and create an account once you're logged in choose settings then access tokens and click create a new token give the token a name and allow a read access to all public gated repositories you can access make sure to copy the token and you also need to accept the Gemma license right search for Gemma 2 and acknowledge the license agreement now back to vertex AI fill in your hugging face access token click deploy wait for a while and make a prediction request alternative ly you can also deploy using Google kubernetes engine make sure you have a gpe autopilot cluster with gpus ready or if you're using standard enable note out of provisioning then copy and apply the following manifest it creates a deployment with the hugging face deep learning containers and a service to expose the deployment it also creates a secret with your hugging phase token because it needs to pull those models from up so that's that you can thousands of models from the hugging face up without leaving the Google Cloud Web console [Music]

Original Description

Model Garden → https://goo.gle/3zXNQ5P Explore AI models in Model Garden → https://goo.gle/3Yn4zJ2 Hugging Face models → https://goo.gle/400RczH You can deploy thousands of AI models from the Hugging Face Hub directly through the Google Cloud web console without setting up complex infrastructure. Vertex AI's Model Garden provides access to many models, including those from Hugging Face, supporting diverse data types and tasks. Watch along as Googler, Wietse Venema, walks through deploying a HuggingFace open model on Google Cloud’s Model Garden. Chapters: 0:00 - Intro 0:19 - Get started with Vertex AI Model Garden 0:55 - Deploying HuggingFace models on Vertex AI 2:27 - Get started today Watch more Google Cloud: Building with Hugging Face → https://goo.gle/BuildWithHuggingFace Subscribe to Google Cloud Tech → https://goo.gle/GoogleCloudTech #HuggingFace #GoogleCloud Speaker: Wietse Venema Products Mentioned: Vertex AI, Model Garden, Open Source, Gemini
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Google Cloud Tech · Google Cloud Tech · 0 of 60

← Previous Next →
1 I’m going for it #GoogleCloudCertified
I’m going for it #GoogleCloudCertified
Google Cloud Tech
2 I had to get #GoogleCloudCertified
I had to get #GoogleCloudCertified
Google Cloud Tech
3 Be better overall at what you do #GoogleCloudCertified
Be better overall at what you do #GoogleCloudCertified
Google Cloud Tech
4 Cloud Monitoring on our radar #Analysis #Uptime
Cloud Monitoring on our radar #Analysis #Uptime
Google Cloud Tech
5 Introduction to Generative AI Studio
Introduction to Generative AI Studio
Google Cloud Tech
6 How to use Github Actions with Google's Workload Identity Federation
How to use Github Actions with Google's Workload Identity Federation
Google Cloud Tech
7 Introduction to Responsible AI
Introduction to Responsible AI
Google Cloud Tech
8 Networking updates and CDMC-certified architecture
Networking updates and CDMC-certified architecture
Google Cloud Tech
9 Create and use a Cloud Storage bucket
Create and use a Cloud Storage bucket
Google Cloud Tech
10 How to digitize text from documents
How to digitize text from documents
Google Cloud Tech
11 Faster analytical queries with AlloyDB
Faster analytical queries with AlloyDB
Google Cloud Tech
12 Next ‘23 sessions and FaaS Wave
Next ‘23 sessions and FaaS Wave
Google Cloud Tech
13 Introduction to Assured Open Source Software
Introduction to Assured Open Source Software
Google Cloud Tech
14 BigQuery Cost Optimization: Storage
BigQuery Cost Optimization: Storage
Google Cloud Tech
15 BigQuery Cost Optimization: Compute
BigQuery Cost Optimization: Compute
Google Cloud Tech
16 BigQuery Cost Optimization: Select Queries
BigQuery Cost Optimization: Select Queries
Google Cloud Tech
17 Remote Field Equipment Management with Manufacturing Data Engine
Remote Field Equipment Management with Manufacturing Data Engine
Google Cloud Tech
18 Supercharging your applications with Cloud SQL Enterprise Plus
Supercharging your applications with Cloud SQL Enterprise Plus
Google Cloud Tech
19 Vector Support on our radar #GenAI
Vector Support on our radar #GenAI
Google Cloud Tech
20 Architecting a blockchain startup with Google Cloud
Architecting a blockchain startup with Google Cloud
Google Cloud Tech
21 Kubernetes and multitasking updates!
Kubernetes and multitasking updates!
Google Cloud Tech
22 GKE: Using Kubernetes Events
GKE: Using Kubernetes Events
Google Cloud Tech
23 How to configure firewall rules for Cloud Composer
How to configure firewall rules for Cloud Composer
Google Cloud Tech
24 Vertex AI Embeddings API + Matching Engine: Grounding LLMs made easy
Vertex AI Embeddings API + Matching Engine: Grounding LLMs made easy
Google Cloud Tech
25 Geospatial analytics on our radar #EarthEngine #BigQuery
Geospatial analytics on our radar #EarthEngine #BigQuery
Google Cloud Tech
26 Ensuring requests are set in Kubernetes
Ensuring requests are set in Kubernetes
Google Cloud Tech
27 Cloud Next 2023, Google research program, and more!
Cloud Next 2023, Google research program, and more!
Google Cloud Tech
28 How to migrate projects between organizations with Resource Manager
How to migrate projects between organizations with Resource Manager
Google Cloud Tech
29 How to run #MySQL in Google Cloud
How to run #MySQL in Google Cloud
Google Cloud Tech
30 #GenerativeAI for enterprises and #Next2023
#GenerativeAI for enterprises and #Next2023
Google Cloud Tech
31 How Google Photos scales to store 4 trillion photos and videos
How Google Photos scales to store 4 trillion photos and videos
Google Cloud Tech
32 Google Cross-Cloud Interconnect (Demo 2)
Google Cross-Cloud Interconnect (Demo 2)
Google Cloud Tech
33 GKE Cost Optimization Golden Signals: Introduction
GKE Cost Optimization Golden Signals: Introduction
Google Cloud Tech
34 GKE Cost Optimization Golden Signals: Workload Rightsizing
GKE Cost Optimization Golden Signals: Workload Rightsizing
Google Cloud Tech
35 GKE Load Balancing: Overview
GKE Load Balancing: Overview
Google Cloud Tech
36 GKE Load Balancing: Best Practices
GKE Load Balancing: Best Practices
Google Cloud Tech
37 Disaster Recovery in GKE
Disaster Recovery in GKE
Google Cloud Tech
38 How to configure IP masquerade agent in GKE Standard clusters
How to configure IP masquerade agent in GKE Standard clusters
Google Cloud Tech
39 Enable and use GKE Control plane logs
Enable and use GKE Control plane logs
Google Cloud Tech
40 Compliance in Australia with Assured Workloads
Compliance in Australia with Assured Workloads
Google Cloud Tech
41 Creating budgets and budget alerts in Google Cloud #FinOps
Creating budgets and budget alerts in Google Cloud #FinOps
Google Cloud Tech
42 Cloud SQL Enterprise Plus on our radar #mySQL
Cloud SQL Enterprise Plus on our radar #mySQL
Google Cloud Tech
43 What's Next for Google Cloud?
What's Next for Google Cloud?
Google Cloud Tech
44 How Loveholidays scaled with Contact Center AI
How Loveholidays scaled with Contact Center AI
Google Cloud Tech
45 What is fleet team management in GKE?
What is fleet team management in GKE?
Google Cloud Tech
46 Troubleshoot VPC Network Peering
Troubleshoot VPC Network Peering
Google Cloud Tech
47 Introduction to DocAI and Contact Center AI
Introduction to DocAI and Contact Center AI
Google Cloud Tech
48 Cloud Run Direct VPC egress explained
Cloud Run Direct VPC egress explained
Google Cloud Tech
49 Database deployment options in GKE
Database deployment options in GKE
Google Cloud Tech
50 Analyze cloud billing data with #BigQuery
Analyze cloud billing data with #BigQuery
Google Cloud Tech
51 Tips to becoming a world-class Prompt Engineer
Tips to becoming a world-class Prompt Engineer
Google Cloud Tech
52 Serverless is simple. Do I need CI/CD?
Serverless is simple. Do I need CI/CD?
Google Cloud Tech
53 Accelerating model deployment with MLOps
Accelerating model deployment with MLOps
Google Cloud Tech
54 How Hawaii's Department of Human Services scaled with CCAI
How Hawaii's Department of Human Services scaled with CCAI
Google Cloud Tech
55 Pricing API on our #Radar
Pricing API on our #Radar
Google Cloud Tech
56 How Recommendations AI for Media can boost customer retention
How Recommendations AI for Media can boost customer retention
Google Cloud Tech
57 Troubleshooting: Node Not Ready Status
Troubleshooting: Node Not Ready Status
Google Cloud Tech
58 One weekend until Cloud Next 2023!
One weekend until Cloud Next 2023!
Google Cloud Tech
59 #GoogleCloudNext starts tomorrow!
#GoogleCloudNext starts tomorrow!
Google Cloud Tech
60 #GoogleCloudNext will be demand!
#GoogleCloudNext will be demand!
Google Cloud Tech

This video demonstrates how to deploy thousands of AI models from the Hugging Face Hub directly through the Google Cloud Web console, using Vertex AI Model Garden and Kubernetes Engine. It covers creating a Hugging Face access token, deploying a model, and making prediction requests.

Key Takeaways
  1. Create a Hugging Face account
  2. Generate an access token
  3. Search for models in Vertex AI Model Garden
  4. Deploy a model using the Google Cloud Web console
  5. Alternatively, deploy using Google Kubernetes Engine
  6. Create a GPE Autopilot cluster with GPUs
  7. Apply the manifest to create a deployment and service
💡 The Hugging Face Hub provides access to thousands of machine learning models, which can be easily deployed and managed using Vertex AI Model Garden and Google Cloud infrastructure.

Related Reads

📰
My LLM drift tracker flagged four regressions this week. All four were wrong.
Learn to track LLM model performance regressions and troubleshoot false positives, crucial for maintaining AI model reliability
Dev.to AI
📰
I built a memory layer that works across Claude, ChatGPT and Cursor — and rendered it as a 3D brain
Build a unified memory layer for multiple AI tools like ChatGPT, Claude, and Cursor to retain context across sessions
Dev.to AI
📰
Put an LLM QA gate in front of your paid deliverable: fail-closed review, one self-heal, no retry loops
Learn to integrate an LLM QA gate for fail-closed review and self-healing in automated compliance reporting using Node and Playwright
Dev.to · Kynth
📰
Jailbreaking de LLMs mediante Prompt Injection
Learn to jailbreak LLMs using prompt injection and engineering techniques to unlock their full potential
Medium · AI

Chapters (4)

Intro
0:19 Get started with Vertex AI Model Garden
0:55 Deploying HuggingFace models on Vertex AI
2:27 Get started today
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →