Testing LLMs Smarter: Multi-Model Experiments & Insights
Key Takeaways
The video demonstrates how to use FloTorch to test and evaluate multiple LLMs, running experiments across different models and prompts, and measuring metrics like context precision, answer relevancy, and response time.
Full Transcript
Okay, I have this bedrock knowledge base I have indexed right with large amount of data. Okay, and that I have also uploaded a ground truth file of 50 questions and answers. I will I want every different combination with different prompt options, right? I can have a zeros prompt, one shot, twoot, threeshot prompt and I attach my prompt file as well. And then I have this system prompt, right? And I have selected four different models. Right? Now remember these could be caching models, routing models, it could be with guardrails, without guardrails. Different options I have selected here from my flow to registry. I have selected two different options for uh nearest neighbor. So when the vector is retrieves, I want top three chunks or top five chunks. I'm using Ragus for my evaluation and I would like model one which is not being used here to test. Right? So I you don't want to double dip you don't want to use evaluation model same as the retrieval model right that is doing inference. So retrieval models are this model 5678 but my evaluation is being done with a larger model in this case it's I'm just naming it model one and my queries would be embedded with Titan embedding uh from bedrock I could use something else as well these are choices I have right so you can go through these right you can select multiple prompts options for uh you can select your knowledge base that you want to select from or kind of just test it out of the box without knowledge is you can use multiple KN&N options like I said and you can choose your model to evaluate with Ragus. If you're not familiar with Ragus and DPL are two standard libraries for LLM evaluation. There are multiple ones actually but these two are pretty commonly used and they are focused on context precision, context recall, right? And then there is maliciousness and all that. Depending on your use case, you can use specific ones but context precision and recall are very commonly used, right? So we are going to do the same thing. We are going to use ragas for context precision and recall. So you'll get results of multiple experiments. So the if you see the number of combinations could be very well 30 40 different combinations in flowch. If you set it up like this and say create project, it will actually run these 30 40 experiments for you using the models you registered with the combinations and we'll do the evaluation with the ground truth file that you provided and report everything on the dashboard like this. Right? And it will tell you what the expected cost is. That's called directional cost. It will tell you how much time it took and the corresponding precision and recall. Answer relevancy is how relevant the answer is. So that's the most important metric in this particular case. So I sorted this table by uh answer 11 C. You'll notice one thing right that inferencing model is same model 7 right but it has different context precision and answer 11 C as well as duration
Original Description
See how FloTorch makes testing and evaluating multiple LLMs easy! In this clip, we explore running experiments across different models and prompts, measuring metrics like context precision, answer relevancy, and response time. Perfect for AI builders looking to optimize LLM performance and reduce costs.”
#flotorch #LLM #AIExperiments #MachineLearning #AIOptimization #RAG #ContextPrecision #LargeLanguageModels
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Data Science Dojo · Data Science Dojo · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Feature Engineering and Predictive Modeling | Data Analytics with R and Azure ML | Community Webinar
Data Science Dojo
Data Exploration and Visualization | Beginning Azure ML | Part 3
Data Science Dojo
Reading External Data Sources | Beginning Azure ML | Part 2
Data Science Dojo
Importing Data, Accessing, & Creating a New Experiment | Beginning Azure ML | Part 1
Data Science Dojo
Casting Columns & Renaming Columns | Beginning Azure ML | Part 4
Data Science Dojo
Scrub Missing Values & Project Columns | Beginning Azure ML | Part 5
Data Science Dojo
Feature Engineering & R Script | Beginning Azure ML | Part 6
Data Science Dojo
Building Your First Model | Beginning Azure ML | Part 7
Data Science Dojo
Run and Fine-Tune Multiple Models | Beginning Azure ML | Part 8
Data Science Dojo
Deploying Your First Predictive Model As a Web Service | Beginning Azure ML | Part 9
Data Science Dojo
Using R API to Obtain Predictions From Your Web Service Beginning Azure ML | Part 10
Data Science Dojo
Using Python API to Obtain Predictions From Your Web Service | Beginning Azure ML | Part 11
Data Science Dojo
Twitter Sentiment Analysis | Natural Language Processing | Community Webinar
Data Science Dojo
Listening to the Melody of the Universe (LIGO Gravitational Waves Presentation) | Community Webinar
Data Science Dojo
David Wechsler on the Impact of Data Science Bootcamp
Data Science Dojo
Andrew Choi on the Impact of Data Science Bootcamp
Data Science Dojo
Microsoft's Software Engineer Shares Her Experience with Data Science Bootcamp
Data Science Dojo
Michael DAndrea on the Impact of Data Science Bootcamp
Data Science Dojo
Data Driven Decision-Making with Data Science Bootcamp: Artem Kopelev's Revelation
Data Science Dojo
Learn the Fundamentals of Data Science: Srinivas Rao's Experience with Data Science Bootcamp
Data Science Dojo
Re-Learning Data Science with Data Science Bootcamp: Analyst's Revelation
Data Science Dojo
Scale R to Big Data with Hadoop & Spark | Community Webinar
Data Science Dojo
Enhancing Skills with Data Science Bootcamp: Sharon Lane-Getaz's Revelation
Data Science Dojo
Ryan DeMartino on the Impact of Data Science Bootcamp
Data Science Dojo
Software Engineer at Microsoft Reveals About His Experience with Data Science Bootcamp
Data Science Dojo
Wade Wimer on the Impact of Data Science Bootcamp
Data Science Dojo
Analyzing Data with Data Science Bootcamp: Hannah Richta's Revelation
Data Science Dojo
Applying Data Science Skills to The Current Role with Bootcamp: Marcos Lacayo's Revelation
Data Science Dojo
Lance Milner on the Impact of Data Science Bootcamp
Data Science Dojo
Deloitte's Data Scientist Revelation: Learning Predictive Analytics with Data Science Bootcamp
Data Science Dojo
Rajesh Patil's Experience at Data Science Bootcamp As an Enterprise Architect
Data Science Dojo
Michael Atlin on the Impact of Data Science Bootcamp
Data Science Dojo
Amina Tariq's In-Person Experience at Data Science Bootcamp
Data Science Dojo
Ceo's Revelation about Data Science Bootcamp
Data Science Dojo
Stephen Miller Describes His Experience at Data Science Dojo's Bootcamp
Data Science Dojo
Kevin Hillaker on the Impact of Data Science Bootcamp
Data Science Dojo
Marko Topalovic's Experience with Data Science Bootcamp
Data Science Dojo
Text Analytics With Python, Cognitive Services & PowerBI | Data Analytics | Community Webinar
Data Science Dojo
Unisys Manager's Revelation: Visualizing Real Time Data with Data Science Bootcamp
Data Science Dojo
Learn Data Mining with Data Science Bootcamp: Ryan LaBrie's Revelation
Data Science Dojo
Vang Xiong on the Impact of Data Science Bootcamp
Data Science Dojo
Data Scientist's Experience at Our Data Science Bootcamp
Data Science Dojo
Alejandro Wolf Yadlin on the Impact of Data Science Bootcamp
Data Science Dojo
Introduction To Titanic Kaggle Competition | Part 1
Data Science Dojo
Learning How to Code in R with Data Science Bootcamp: Priscilla Mannuel's Revelation
Data Science Dojo
Andrew Berman On Why Data Science Bootcamp Is Better Fit for Him
Data Science Dojo
How To Do Titanic Kaggle Competition in R | Part 3.1
Data Science Dojo
How to do the Titanic Kaggle competition in R | Part 3.1
Data Science Dojo
Delve Deeper into Data Science with Data Science Bootcamp
Data Science Dojo
Bank of America Data Scientist Reveals His Experience of Data Science Bootcamp
Data Science Dojo
Shaena Montanari on the Impact of Data Science Bootcamp
Data Science Dojo
Types of Sampling | Introduction to Data Mining | Part 12
Data Science Dojo
Sampling for Data Selection | Introduction to Data Mining | Part 11
Data Science Dojo
Data Aggregation | Introduction to Data Mining | Part 10
Data Science Dojo
Data Cleaning | Introduction to Data Mining | Part 9
Data Science Dojo
Missing & Duplicated Data | Introduction to Data Mining | Part 8
Data Science Dojo
Data Noise | Introduction to Data Mining | Part 7
Data Science Dojo
Graph and Ordered Data | Introduction to Data Mining | Part 5
Data Science Dojo
Document Data & Transaction Data | Introduction to Data Mining | Part 4
Data Science Dojo
Data Quality | Introduction to Data Mining | Part 6
Data Science Dojo
More on: LLM Engineering
View skill →Related Reads
📰
📰
📰
📰
I ran a 110B LLM on 16GB of RAM. Here's the equation that predicts any model's speed on your machine
Dev.to · Federico Sciuca
A Fidelity-First Workflow for Editing GPT-Generated Text
Dev.to · Bisrat
Never Let the Model Pick the Tenant ID: Securing an LLM Agent in Go
Dev.to · Jules Robineau
The Research Assistant in the Room
Dev.to · Thomas Lee
🎓
Tutor Explanation
DeepCamp AI