Optimizing agentic AI workflows: Metrics-driven evaluation with W&B Weave and NVIDIA - FC London '25

Weights & Biases · Intermediate ·🤖 AI Agents & Automation ·7mo ago

Key Takeaways

This video demonstrates the use of W&B Weave and NVIDIA NeMo Agent Toolkit for optimizing agentic AI workflows, with a focus on metrics-driven evaluation and multi-step reasoning and tool use. The session covers the implementation of alert triage systems and the evaluation of AI pipelines using various metrics and tools.

Full Transcript

Hi everyone. Uh bearing in mind that I'm the only thing between uh you and drinks. Let's get this going. If you cannot see my face uh directly, you can see my face in the slide. So I'm Retita. I work for Nvidia as a solution architect. Most of what I do connects to evaluations. I'm going to talk about agentic systems and evaluation, but feel free to find me outside of this room and ask me about what I think about the benchmarks that we currently use. My background is um space engineering. So also don't ask me um don't tell me that it's not rocket science because I really hate it. Everything is complex. Okay, especially re uh generative AI. So let's start from the beginning which is um the latest in artificial intelligence. I think you may have seen this or a similar slide today before but there have been many phases to generative AI. First we we had the LLM boom uh with CHP in 2022 and then the adoption of LLMs. Then we were talking about rack systems and we're still talking about Rex systems if my job has anything to say about it. And finally now we're going to be talking way more about agents. Um, it's not that we've solved the LLMs in rag system problems, but now we decided that beyond just connecting an LLM to a vector database, maybe we want to connect an LLM to a bunch of other different stuff with different protocols, which makes everything much more complex and makes all of our jobs much harder, but also hopefully our lives much better in the future. Okay. So, how do AI agents work? At a high level, what an agent does is use reasoning and planning to solve multi-step problems. And the way it works is through a few core components. One of them is what we call thinking. But of course, it's thinking like a machine can think. And that's mostly enabled by reasoning capabilities. So, the capabilities of an LLM to go through a chain of thought, stepbystep problem solving. Then we have of course prompting which starts from a user's goal or question and interpreting intent. And then we have what I mentioned before which is the tool use which is calling different APIs, calling the web, calling any kind of dashboards, computer use and memory of course which we may also want to to implement. And finally, we have collaboration as well, which means that we may want our LLMs to speak to each other. And again, speak is a very loosely used word here. So this is a good conceptual understanding of what agents can be, which basically means that we're taking an LLM and we're adding a lot of different tools and software on top of it basically to gain observability, monitoring, evaluation capabilities and for it to be um essentially an end toend scenario that users can use to develop their applications. So let's talk a bit more about the software that's necessary for agentic implementations. Um as as I'm representing Nvidia, I'm going to talk a bit about what Nvidia has for agentic systems and we tend to split this into three different components. The first one are the Nvidia Nims which are essentially models that are containerized in a micros service. um we have a lot of open-source models and also proprietary models uh in our NIMS. Then we have the AI blueprints and the blueprints and I'm going to do a demo by the end are our reference architectures for different AI pipelines. So for instance research agents, customer service agents etc. And by reference architecture I mean that it's an implementation using particular libraries and models but you can go and modularly change them out. So it's just a reference for developers to then improve upon. And finally, we have Nemo, which is a software suite composed of both SDKs and microservices whose goal is to basically serve all of the components of an AI pipeline from the data curation with Nemo curator to customization which is a post-raining library to evaluation and guardrails. Evaluation of course is to evaluate models with academic benchmarks or with user brought custom benchmarks or also with LLM as a judge and finally guard rails are to control basically the inputs and outputs of the LLM to whatever degree we want. Connecting all of these is our most recent agentic library which is what I'm going to be talking about here which is the Nemo agent toolkit. So Nemo agent toolkit is open source. So you're free to check it on GitHub or on our own website build.envidia.com. And the goal is to basically make aic implementation less complex. Um it includes everything that I've mentioned before from the monitoring tooling observability and also evaluator and guard rails. Uh and it is modular, flexible and also is integrated with very different uh components. One of them is weave. So the integration with uh weight weights and biases weave is what I'm talking about here today. Let's talk about it. Um so the weights and biases team has developed weep to make it easy for organizations building genai um end to end pipelines to monitor and evaluate and essentially have a continuous um eye into what's happening into an AI pipeline. And so the capabilities include tracing running and tracking evaluations and finally it's framework and LLM agnostics. So you can integrate any model with it. Uh and one of the reasons um I'm here today is because you can integrate it with an emo agent toolkit. So uh here is very small so I apologize for that. the typical workflow to build an agent. But in here, basically we just show uh all of the steps of monitoring and observability both on an automated side and on a manual side that you can also implement. And what Nemo agent toolkit does is it produces those traces. So it runs the agent logic. It collects events and defines exactly the metrics that you want to measure. And what we've does complementary is uh handle the logging, the visualization and experiment management. So basically we have Nemo agent toolkit as an instrumentation plus runtime for agents that you may want to implement and we've adds the observability and visualization layer. So here is an example of uh the integration you may have. Again maybe a bit small for you but essentially this is one of the dashboards that uh allows to you for you to check the evaluation uh for a particular Gentic application. So let's look at some examples. First one is one of the blueprints that we have uh internally and also uh on GitHub. So it's open source. It's called AIQ. Um this Agentic application is a research agent. So what it does is deep research and you may be familiar with this from Gemini deep research or JPT. Deep research. Deep research is basically um a name used everywhere for when you want to build uh when you want to do a lot of different searches and collate information into a very extensive let's say report and so this is an a blueprint which I've mentioned before is a reference architecture so you may when you use it you are free to basically change the LLMs change any kind of libraries and see if it makes more sense for your application. So, uh, the AIQ blueprint searches many different sources and the web uh, to synthesize hours of research into a concise report as I've mentioned and is currently the top ranked open agent on deep research bench in hugging face. So, feel free to try it out. And the one I'm going to speak about with my limited time mostly is the alert triage agent which is also one of our blueprints. And this one we have integrated with weave. We have a blog post about it that I recommend. It's very interesting. And the goal here was to automate uh the monitor and action of alerts in largecale environments. So uh this saves a lot of time in manual checks when you have a large scale implementation of a software um scenario where you have different fail points that you need to constantly check for problems um or any kind of corner or edge cases. So as an alert agent this requires a lot of diagnostics for hardware telemetry and service health and you can see them here also on this diagram if you are able to see a very small screen and um with Nemo agent toolkit you can implement and evaluate uh any custom metrics for this kind of application. So um here is also an example of the tracing and evaluating that you can do with weave here in these plots. But what I really want to get to is the demo which uh I'm not going to be speaking but there's going to be sound. This is about four minutes so please bear with me and it's going to be showing how the implementation of the alert triaging system works with weave. I'll demonstrate how to evaluate a data set of alerts using the alert triage agent example from the NVIDIA Nemo agent toolkit repository. Here's the configuration file for the alert triage agent workflow. I'll run two valuations, one using the llama 3.18b model and the other using llama 3.370B. We'll use weights and biases view to analyze the runs and view the evaluation results. Let's start the first evaluation with the lama 3.18b model. In the evaluation logs, you'll see the URL for the V dashboard. The first run is complete. Now launching the second evaluation using Llama 3.370B. Both runs are complete. Let's take a look at the results in the Weave dashboard. By default, the VE dashboard displays the trace logs for each run. I'll switch to the eval tab in the sidebar to view the evaluation results. There are several runs here from earlier experiments. I'll compare these two lama 3.1 and lama 3.3 results. We are evaluating two metrics classification accuracy a custom evaluator builds for the alert triage agent and rag accuracy which is the built-in raggas evaluator. Across both metrics we see an improvement in accuracy. In addition to accuracy, the summary tab displays key performance metrics such as the LLM latency, the P95 outflow runtime. This represents how long it takes to process a single entry, total runtime, the overall time to process all data set entries. As we're running the data set entries in parallel, we can see that the total runtime is close to the P95 outflow runtime. Total tokens. This is the token usage across the entire data set. Now let's go to the data set results tab. Here you can pick and plots any metric across different runs. You can also inspect details per entry results in the table below. In addition to these evaluation resource, you can view all the tools and MLM calls made by the workflow via the traces tab. You can pick an evaluation run here. To view the individual entries, you select the table of child calls. In this table, you can select a data set entry and view a detailed breakdown of all the calls made by the workflow. Uh hopefully it was useful. Uh essentially I realize I didn't speak a lot about evaluation because in this particular case you may choose your own evaluation metrics. The demo included classification accuracy and raggas metrics which are of course included in Nemo evaluator which is integrated into Neoentic toolkit. There's a lot of different metrics and benchmarks in Nemo in Nemo Evaluator as well um including tool calling metrics, agentic uh benchmarks and also the regular academic benchmarks that you know like a25 I'm happy to talk about that after this talk but uh the key takeaways are essentially that Nemo agentic toolkit allows you to do a lot of the heavy lifting on creating agents especially regarding evaluation profiling and all of the tool calling and uh integrating with weights and biases we've allows you to have those dashboards and observability that are always necessary when implementing an end toend AI solution and uh if you want to know more because this is is this is only a 20inut conversation everything that we have presented here is open source so you're free to check GitHub for that. Uh, of course, models have to be deployed. Uh, for the AI deep research blueprint, you have here the QR code for the Nemo agent toolkit as well and for weave in the end, although I'm sure you have access to it very easily. Anything else you can find on the build.envidia.com website, which is a fairly fun website also to try out different LLMs. And that's it for my talk. I think right on time. So, thank you.

Original Description

In this session from Fully Connected London, Rita Fernandes Neves, Sr. Solutions Architect at NVIDIA, explores how to build, debug, and optimize agentic workflows—where large language models drive multi-step reasoning and tool use. She demonstrates how W&B Weave and the NVIDIA NeMo Agent Toolkit enable transparent, metrics-focused development of agentic applications. Rita showcases practical strategies for workflow instrumentation, evaluation, and visualization across structured operations and open-ended research assistants. Learn how YAML-driven configuration, automated and human-in-the-loop evaluation, and intuitive dash boarding help teams benchmark models, diagnose bottlenecks, and iteratively refine AI agents for reliability and performance.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Weights & Biases · Weights & Biases · 0 of 60

← Previous Next →
1 0. What is machine learning?
0. What is machine learning?
Weights & Biases
2 1. Build Your First Machine Learning Model
1. Build Your First Machine Learning Model
Weights & Biases
3 Intro to ML: Course Overview
Intro to ML: Course Overview
Weights & Biases
4 2. Multi-Layer Perceptrons
2. Multi-Layer Perceptrons
Weights & Biases
5 3. Convolutional Neural Networks
3. Convolutional Neural Networks
Weights & Biases
6 Weights & Biases at OpenAI
Weights & Biases at OpenAI
Weights & Biases
7 Why Experiment Tracking is Crucial to OpenAI
Why Experiment Tracking is Crucial to OpenAI
Weights & Biases
8 4. Autoencoders
4. Autoencoders
Weights & Biases
9 5. Sentiment Analysis
5. Sentiment Analysis
Weights & Biases
10 6. Recurrent Neural Networks [RNNs]
6. Recurrent Neural Networks [RNNs]
Weights & Biases
11 7. Text Generation using LSTMs and GRUs
7. Text Generation using LSTMs and GRUs
Weights & Biases
12 8. Text Classification Using Convolutional Neural Networks
8. Text Classification Using Convolutional Neural Networks
Weights & Biases
13 9. Hybrid LSTMs [Long Short-Term Memory]
9. Hybrid LSTMs [Long Short-Term Memory]
Weights & Biases
14 Toyota Research Institute on Experiment Tracking with Weights & Biases
Toyota Research Institute on Experiment Tracking with Weights & Biases
Weights & Biases
15 Weights and Biases - Developer Tools for Deep Learning
Weights and Biases - Developer Tools for Deep Learning
Weights & Biases
16 Introducing Weights & Biases
Introducing Weights & Biases
Weights & Biases
17 10. Seq2Seq Models
10. Seq2Seq Models
Weights & Biases
18 11. Transfer Learning for Domain-Specific Image Classification with Small Datasets
11. Transfer Learning for Domain-Specific Image Classification with Small Datasets
Weights & Biases
19 12. One-shot learning for teaching neural networks to classify objects never seen before
12. One-shot learning for teaching neural networks to classify objects never seen before
Weights & Biases
20 13. Speech Recognition with Convolutional Neural Networks in Keras/TensorFlow
13. Speech Recognition with Convolutional Neural Networks in Keras/TensorFlow
Weights & Biases
21 14. Data Augmentation | Keras
14. Data Augmentation | Keras
Weights & Biases
22 15. Batch Size and Learning Rate in CNNs
15. Batch Size and Learning Rate in CNNs
Weights & Biases
23 Applied Deep Learning Fellowship Overview and Project Selection with Josh Tobin (2019)
Applied Deep Learning Fellowship Overview and Project Selection with Josh Tobin (2019)
Weights & Biases
24 Grading Rubric for AI Applications with Sergey Karayev  (2019)
Grading Rubric for AI Applications with Sergey Karayev (2019)
Weights & Biases
25 16. Video Frame Prediction using CNNs and LSTMs (2019)
16. Video Frame Prediction using CNNs and LSTMs (2019)
Weights & Biases
26 Image to LaTeX - Applied Deep Learning Fellowship (2019)
Image to LaTeX - Applied Deep Learning Fellowship (2019)
Weights & Biases
27 17.  Build and Deploy an Emotion Classifier (2019)
17. Build and Deploy an Emotion Classifier (2019)
Weights & Biases
28 Applied Deep Learning - Data Management with Josh Tobin (2019)
Applied Deep Learning - Data Management with Josh Tobin (2019)
Weights & Biases
29 Snorkel: Programming Training Data with Paroma Varma of Stanford University (2019)
Snorkel: Programming Training Data with Paroma Varma of Stanford University (2019)
Weights & Biases
30 Applied Deep Learning - Troubleshooting and Debugging with Josh Tobin (2019)
Applied Deep Learning - Troubleshooting and Debugging with Josh Tobin (2019)
Weights & Biases
31 Troubleshooting and Iterating ML Models with Lee Redden (2019)
Troubleshooting and Iterating ML Models with Lee Redden (2019)
Weights & Biases
32 Designing a Machine Learning Project with Neal Khosla (2019)
Designing a Machine Learning Project with Neal Khosla (2019)
Weights & Biases
33 Lukas Beiwald on ML Tools and Experiment Management (2019)
Lukas Beiwald on ML Tools and Experiment Management (2019)
Weights & Biases
34 Building Machine Learning Teams with Josh Tobin (2019)
Building Machine Learning Teams with Josh Tobin (2019)
Weights & Biases
35 Pieter Abeel on Potential Deep Learning Research Directions  (2019)
Pieter Abeel on Potential Deep Learning Research Directions (2019)
Weights & Biases
36 Testing and Deployment of Deep Learning Models with Josh Tobin (2019)
Testing and Deployment of Deep Learning Models with Josh Tobin (2019)
Weights & Biases
37 Five Lessons for Team-Oriented Research with Peter Welder (2019)
Five Lessons for Team-Oriented Research with Peter Welder (2019)
Weights & Biases
38 Applied Deep Learning - Rosanne Liu on AI Research (2019)
Applied Deep Learning - Rosanne Liu on AI Research (2019)
Weights & Biases
39 Making the Mid-career Leap from Urban Design to Deep Learning/Data Science
Making the Mid-career Leap from Urban Design to Deep Learning/Data Science
Weights & Biases
40 Organizing ML projects — W&B walkthrough (2020)
Organizing ML projects — W&B walkthrough (2020)
Weights & Biases
41 Brandon Rohrer — Machine Learning in Production for Robots
Brandon Rohrer — Machine Learning in Production for Robots
Weights & Biases
42 Nicolas Koumchatzky — Machine Learning in Production for Self-Driving Cars
Nicolas Koumchatzky — Machine Learning in Production for Self-Driving Cars
Weights & Biases
43 My experiments with Reinforcement Learning with Jariullah Safi
My experiments with Reinforcement Learning with Jariullah Safi
Weights & Biases
44 Applications of Machine Learning to COVID-19 Research with Isaac Godfried
Applications of Machine Learning to COVID-19 Research with Isaac Godfried
Weights & Biases
45 Testing Machine Learning Models with Eric Schles
Testing Machine Learning Models with Eric Schles
Weights & Biases
46 How Linear Algebra is not like Algebra with Charles Frye
How Linear Algebra is not like Algebra with Charles Frye
Weights & Biases
47 Predicting Protein Structures using Deep Learning with Jonathan King
Predicting Protein Structures using Deep Learning with Jonathan King
Weights & Biases
48 Rachael Tatman — Conversational AI and Linguistics
Rachael Tatman — Conversational AI and Linguistics
Weights & Biases
49 Reformer by Han Lee
Reformer by Han Lee
Weights & Biases
50 Sequence Models with Pujaa Rajan
Sequence Models with Pujaa Rajan
Weights & Biases
51 GitHub Actions & Machine Learning Workflows with Hamel Husain
GitHub Actions & Machine Learning Workflows with Hamel Husain
Weights & Biases
52 Look Mom, No Indices! Vector Calculus with the Fréchet Derivative by Charles Frye
Look Mom, No Indices! Vector Calculus with the Fréchet Derivative by Charles Frye
Weights & Biases
53 Jack Clark — Building Trustworthy AI Systems
Jack Clark — Building Trustworthy AI Systems
Weights & Biases
54 Surprising Utility of Surprise: Why ML Uses Negative Log Probabilities - Charles Frye
Surprising Utility of Surprise: Why ML Uses Negative Log Probabilities - Charles Frye
Weights & Biases
55 Track your machine learning experiments locally, with W&B Local - Chris Van Pelt
Track your machine learning experiments locally, with W&B Local - Chris Van Pelt
Weights & Biases
56 Antipatterns in open source research code with Jariullah Safi
Antipatterns in open source research code with Jariullah Safi
Weights & Biases
57 Attention for time series forecasting & COVID predictions - Isaac Godfried
Attention for time series forecasting & COVID predictions - Isaac Godfried
Weights & Biases
58 Made with ML - Goku Mohandas
Made with ML - Goku Mohandas
Weights & Biases
59 Angela & Danielle — Designing ML Models for Millions of Consumer Robots
Angela & Danielle — Designing ML Models for Millions of Consumer Robots
Weights & Biases
60 Deep Learning Salon by Weights & Biases
Deep Learning Salon by Weights & Biases
Weights & Biases

This video teaches how to optimize agentic AI workflows using W&B Weave and NVIDIA NeMo Agent Toolkit, with a focus on metrics-driven evaluation and multi-step reasoning and tool use. The session provides a comprehensive overview of the tools and techniques used in agentic AI workflows, including the implementation of alert triage systems and the evaluation of AI pipelines. By watching this video, viewers can learn how to build, debug, and optimize agentic workflows, and how to use various metri

Key Takeaways
  1. Run two valuations using the alert triage agent example from the NVIDIA Nemo agent toolkit repository
  2. Configure the alert triage agent workflow
  3. Launch the second evaluation using Llama 3.370B
  4. Compare the results in the Weave dashboard
  5. Use W&B Weave for monitoring and evaluation
  6. Implement NVIDIA NeMo Agent Toolkit for AI pipeline management
💡 The use of metrics-driven evaluation and multi-step reasoning and tool use is crucial for optimizing agentic AI workflows, and tools like W&B Weave and NVIDIA NeMo Agent Toolkit provide a comprehensive framework for building, debugging, and optimizing these workflows.

Related Reads

📰
The Fall of Static Audits: Analyzing AI-Driven Supply Chain Risk Management
Learn how AI-driven supply chain risk management can help mitigate disruptions and minimize losses, making traditional static audits obsolete.
Dev.to AI
📰
Building a multilingual voice agent: lessons from FR/EN/DE: field controls that hold
Learn how to build a multilingual voice agent by applying lessons from French, English, and German language support
Dev.to AI
📰
How I run an AI agent 24/7 on a Raspberry Pi (and don't lose its memory)
Run an AI agent 24/7 on a Raspberry Pi without losing its memory, a cost-effective and secure solution
Dev.to AI
📰
How MADDPG Combines Deep Learning with Multi-Agent Strategies
Learn how MADDPG combines deep learning with multi-agent strategies to enable AI agents to learn, collaborate, and compete in complex environments
Medium · Deep Learning
Up next
Best AI Agent Community to Accelerate Your Learning of AI (James Dooley Chats with Julian Goldie)
James Dooley
Watch →