Creating Embeddings With Closed-Source Model in Weaviate | Vector Databases for Beginners | Part 8
Key Takeaways
Using OpenAI's API with Weaviate to generate vectors automatically and connecting Weaviate Cloud with API for closed-source model embedding creation
Full Transcript
Now I wanted to go through a different example where we use OpenAI. This is also one where you can use cohhere. Um and this is where we're going to send our data through their API since it's a closed source model in order to generate the embeddings. So um I did the same stuff up here. We have the same random but not the same random a different random data set of 100 objects. Um, and now we need to connect to the cloud again here. So, let me go back over here, click the connect button for the right sandbox and copy our URL and also copy our API key. And I've set my OpenAI API key as an environment variable in Google Collab because that's one I don't want you to see. So, [laughter] um, but you can use it in Collab like this or use, um, any other library in Python. So, now we need to connect to the cloud again. So, I'm going to copy this and paste it here. And there's actually one other thing we need to include here, which is a header. So, if we go over to Weeve's documentation and model provider integrations, you'll see all of our different model providers that you can use for vector embeddings slash um generative modules. So, OpenAI has text embeddings. That's the one we want to use. We'll talk about generative modules in the third webinar. You should also join that one. Um, and if we scroll down, we'll see here this will be our header that we need to include in the connect. So, I'm just going to copy this here. Go back over here. Paste it in there. And I named this slightly differently. So, let's change that. Um, forgot to run this, but here's our different data set of 100 objects. And hopefully when I run this, we're going to get a is ready. Nice. True. Awesome. So, [sighs] as you can see, I updated the headers. It's the same collection. Super updatable. Very nice. And now we're going to create a new collection. This one's going to be called OpenAI vectors collection. Um, again, doing the same thing if it exists. Delete it. Now, our vectorzer config is going to be different again. So, we're going to do the same thing. Pull up my notes over here. You can find this in the documentation. Um yeah. Okay. Vectorzer config is going to be WC dot configure dot vectorzer dot text to vec open ai parenthesis. And that's it. So hopefully we run this and again we've defined properties. So we have a text and a title and basically behind the scenes what we will do is any data object that you import it will automatically vectorize it using this config. So we run this. We're going to create the collection with those settings and then we can insert our objects again here. This time we're not including the vector property because we is going to do that automatically. So let's do that. It's taking a little bit longer because again it has to vectorize it. And then if we go to our cloud again and then if we go to the explorer tool again and choose our sandbox. Oh, we have a new collection. If we click on that we are seeing data. Awesome. Okay. And we can do the same little searchy thing. This time instead of near vector, we're going to use near text because we have specified a vector. So we can just actually add the text directly. We don't have to vectorize it beforehand. And we get results. and different results cuz it's a different data set. But that's why
Original Description
In this part 7, we move from open-source embedding models to a closed-source model — using OpenAI’s API with Weaviate to generate vectors automatically.
In this section, we're going to go over:
- Using a closed-source model (OpenAI) to create embeddings via API
- Connecting Weaviate Cloud with API keys and provider headers
- Setting up a collection that vectorizes data automatically using text2vec-openai
- Inserting raw data without manually creating embeddings
- Running similarity search using nearText instead of a pre-generated vector
Closed-source models handle vector creation for you — you send text, and the embeddings come back ready to use.
#OpenAI #Weaviate #ClosedSourceModels #APIIntegration #Text2VecOpenAI #SimilaritySearch
Learn data science, AI, and machine learning through our hands-on training programs: https://www.youtube.com/@Datasciencedojo/courses
Check our community webinars in this playlist: https://www.youtube.com/playlist?list=PL8eNk_zTBST-EBv2LDSW9Wx_V4Gy5OPFT
Check our latest Future of Data and AI Conference: https://www.youtube.com/playlist?list=PL8eNk_zTBST9Wkc6-bczfbClBbSKnT2nI
Subscribe to our newsletter for data science content & infographics: https://datasciencedojo.com/newsletter/
Love podcasts? Check out our Future of Data and AI Podcast with industry-expert guests: https://www.youtube.com/playlist?list=PL8eNk_zTBST_jMlmiokwBVfS_BqbAt0z2
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Data Science Dojo · Data Science Dojo · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Feature Engineering and Predictive Modeling | Data Analytics with R and Azure ML | Community Webinar
Data Science Dojo
Data Exploration and Visualization | Beginning Azure ML | Part 3
Data Science Dojo
Reading External Data Sources | Beginning Azure ML | Part 2
Data Science Dojo
Importing Data, Accessing, & Creating a New Experiment | Beginning Azure ML | Part 1
Data Science Dojo
Casting Columns & Renaming Columns | Beginning Azure ML | Part 4
Data Science Dojo
Scrub Missing Values & Project Columns | Beginning Azure ML | Part 5
Data Science Dojo
Feature Engineering & R Script | Beginning Azure ML | Part 6
Data Science Dojo
Building Your First Model | Beginning Azure ML | Part 7
Data Science Dojo
Run and Fine-Tune Multiple Models | Beginning Azure ML | Part 8
Data Science Dojo
Deploying Your First Predictive Model As a Web Service | Beginning Azure ML | Part 9
Data Science Dojo
Using R API to Obtain Predictions From Your Web Service Beginning Azure ML | Part 10
Data Science Dojo
Using Python API to Obtain Predictions From Your Web Service | Beginning Azure ML | Part 11
Data Science Dojo
Twitter Sentiment Analysis | Natural Language Processing | Community Webinar
Data Science Dojo
Listening to the Melody of the Universe (LIGO Gravitational Waves Presentation) | Community Webinar
Data Science Dojo
David Wechsler on the Impact of Data Science Bootcamp
Data Science Dojo
Andrew Choi on the Impact of Data Science Bootcamp
Data Science Dojo
Microsoft's Software Engineer Shares Her Experience with Data Science Bootcamp
Data Science Dojo
Michael DAndrea on the Impact of Data Science Bootcamp
Data Science Dojo
Data Driven Decision-Making with Data Science Bootcamp: Artem Kopelev's Revelation
Data Science Dojo
Learn the Fundamentals of Data Science: Srinivas Rao's Experience with Data Science Bootcamp
Data Science Dojo
Re-Learning Data Science with Data Science Bootcamp: Analyst's Revelation
Data Science Dojo
Scale R to Big Data with Hadoop & Spark | Community Webinar
Data Science Dojo
Enhancing Skills with Data Science Bootcamp: Sharon Lane-Getaz's Revelation
Data Science Dojo
Ryan DeMartino on the Impact of Data Science Bootcamp
Data Science Dojo
Software Engineer at Microsoft Reveals About His Experience with Data Science Bootcamp
Data Science Dojo
Wade Wimer on the Impact of Data Science Bootcamp
Data Science Dojo
Analyzing Data with Data Science Bootcamp: Hannah Richta's Revelation
Data Science Dojo
Applying Data Science Skills to The Current Role with Bootcamp: Marcos Lacayo's Revelation
Data Science Dojo
Lance Milner on the Impact of Data Science Bootcamp
Data Science Dojo
Deloitte's Data Scientist Revelation: Learning Predictive Analytics with Data Science Bootcamp
Data Science Dojo
Rajesh Patil's Experience at Data Science Bootcamp As an Enterprise Architect
Data Science Dojo
Michael Atlin on the Impact of Data Science Bootcamp
Data Science Dojo
Amina Tariq's In-Person Experience at Data Science Bootcamp
Data Science Dojo
Ceo's Revelation about Data Science Bootcamp
Data Science Dojo
Stephen Miller Describes His Experience at Data Science Dojo's Bootcamp
Data Science Dojo
Kevin Hillaker on the Impact of Data Science Bootcamp
Data Science Dojo
Marko Topalovic's Experience with Data Science Bootcamp
Data Science Dojo
Text Analytics With Python, Cognitive Services & PowerBI | Data Analytics | Community Webinar
Data Science Dojo
Unisys Manager's Revelation: Visualizing Real Time Data with Data Science Bootcamp
Data Science Dojo
Learn Data Mining with Data Science Bootcamp: Ryan LaBrie's Revelation
Data Science Dojo
Vang Xiong on the Impact of Data Science Bootcamp
Data Science Dojo
Data Scientist's Experience at Our Data Science Bootcamp
Data Science Dojo
Alejandro Wolf Yadlin on the Impact of Data Science Bootcamp
Data Science Dojo
Introduction To Titanic Kaggle Competition | Part 1
Data Science Dojo
Learning How to Code in R with Data Science Bootcamp: Priscilla Mannuel's Revelation
Data Science Dojo
Andrew Berman On Why Data Science Bootcamp Is Better Fit for Him
Data Science Dojo
How To Do Titanic Kaggle Competition in R | Part 3.1
Data Science Dojo
How to do the Titanic Kaggle competition in R | Part 3.1
Data Science Dojo
Delve Deeper into Data Science with Data Science Bootcamp
Data Science Dojo
Bank of America Data Scientist Reveals His Experience of Data Science Bootcamp
Data Science Dojo
Shaena Montanari on the Impact of Data Science Bootcamp
Data Science Dojo
Types of Sampling | Introduction to Data Mining | Part 12
Data Science Dojo
Sampling for Data Selection | Introduction to Data Mining | Part 11
Data Science Dojo
Data Aggregation | Introduction to Data Mining | Part 10
Data Science Dojo
Data Cleaning | Introduction to Data Mining | Part 9
Data Science Dojo
Missing & Duplicated Data | Introduction to Data Mining | Part 8
Data Science Dojo
Data Noise | Introduction to Data Mining | Part 7
Data Science Dojo
Graph and Ordered Data | Introduction to Data Mining | Part 5
Data Science Dojo
Document Data & Transaction Data | Introduction to Data Mining | Part 4
Data Science Dojo
Data Quality | Introduction to Data Mining | Part 6
Data Science Dojo
More on: RAG Basics
View skill →Related Reads
📰
📰
📰
📰
Treat Retrieved Content as Data, Not Instructions
Dev.to AI
How to Evaluate Production RAG: Keyword, Vector, SQL, and Hybrid Retrieval
Dev.to · Anya Summers
RAG Database Design: SQL, Full-Text Search, Vector Search, and Context Retrieval
Dev.to · puffball1567
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to · Imus
🎓
Tutor Explanation
DeepCamp AI