Building agentic AI applications with W&B Weave
Key Takeaways
Building agentic AI applications using W&B Weave is demonstrated for autonomous AI systems
Full Transcript
[Music] Hi, I'm Russ from Weights and Biases and today I'll be discussing building agentic AI applications using W andBW weave. Before we go any further, let's define AI agents. Agents are autonomous AI systems, meaning that they act on their own without human direction or oversight, allowing them to function entirely independently. They plan and execute multi-step tasks in a loop. While often guided by a plan generated when launched, tasks are often initiated and executed in a cyclical rather than linear manner depending on the job that needs to get done. Agents use memory and tools while they operate to expand their capabilities and enhance their ability to perform complex jobs autonomously. And finally, agents understand their end goal and continue to work until it's been achieved. On the right here, we have a simple highle chart depicting an agentic workflow. An end user interacts with the agent perhaps by asking a question. The agent considers the question and generates a plan of action. The plan involves multiple steps and continues as needed. And finally, when the job is done, the agent concludes its work. Now that we know what an agent is, let's take a look at an agent in action. This is a financial analyst dashboard where we can quickly check stock quotes and market performance. We'll avoid singling out any stocks for the sake of the demo and focus on the NASDAQ composite. So, I just type into the ticker and hit return. Looks like somewhat of a roller coaster. And as we expand out into the 5day and one month views, we're still seeing a lot of volatility. This is a perfect opportunity to use our financial research agent to shed some light on how investors should proceed given the behavior of the market over the past 30 days. Let's ask which sectors within the NASDAQ have shown the most resilience over the past 30 days. After we hit return, we see our agent getting to work and keeping us apprised of every step in the workflow here in the chat window. And what does that workflow look like? Well, it's a bit more complicated than the simplified version we saw on the previous slide. On the left, we see our AI research assistant interface from our analyst dashboard. We provide a research topic and our agent sprints into action. The agent begins by generating a plan. The planner agent creates a list of five to 15 search engine queries to conduct the online research on the topic at hand. After the web content has been collected using the search queries, a financial writer agent summarizes the retrieved content with help from a fundamentals analyst agent and a risk analyst agent. And to wrap things up, a verification agent reviews and determines the accuracy of the summarized report. And once the agent has generated the report, it's ready for viewing in the dashboard interface. Now, let's talk about weave. You saw the agent at work and you saw the workflow. Now let me show you how weave captures all the important data, the inputs, outputs, metadata and code while you develop, evaluate and optimize your AI agentic applications. This trace at the top is the financial research request we just submitted a moment ago. And I can scroll all the way over to the right and you can see some of the aggregated data. total tokens, total cost, latency, and trace size. And when I click on the trace, you see the comprehensive trace tree and all of the input and output data for each agent in each step of the workflow. This comes in especially handy for drilling down into the LLM API calls, so I can see everything that was sent, including the prompt and configuration settings and everything that was received in the output. Here we can see those searches that were generated for assignment to the financial search agents. The trace tree shows you the agent step-by-step operation. The default view shows you all the tasks completed by each agent and you can use this slider at the bottom to scroll through the timeline of task execution. We also have a couple of other trace views that are incredibly useful for examining agents and agentic workflows. This is the code composition view. It organizes traces by grouping and nesting traces within the agent by which they were called. This allows us to quickly visualize how the workflow plan has been executed, which tasks are most closely related to one another and take account of the aggregated metrics for each task in each group, including total number of tasks executed and the total latency. The flame view is an interactive chart that allows us to not only visualize the agentic workflow, but effectively navigate through it and drill down on any task of interest. This layered view allows us to quickly identify concurrent activities, bottlenecks, and dependencies within our workflow, providing deep insights into orchestration, handoffs, performance, and efficiency. So now we've learned that agents like our financial research analysts that may appear somewhat simple on the surface are actually fairly complex. It's not a ton of work to quickly build out a multi-step workflow consisting of LOM calls tied to one another. But building an agent that can be confidently deployed into production is not easy. Why? Agents are prone to make mistakes that humans wouldn't make. This is largely because LMS, the underlying engine for our agents, are non-deterministic. This results in unpredictability. If asked the exact same question a 100 times, they can never be depended on to consistently return the exact same answer. And finally, agents in agentic workflows are complex and opaque. It's challenging to understand exactly what happens in each task, how long it takes, and how it relates to other tasks. But this knowledge is critical when developing a productionready application. And this is where Weave comes in. Making agents easy or at least easier to productionize. By adding just a single line of code, we've tracks and organizes all of your agent inputs, outputs, code, and metadata, allowing you to rapidly iterate. We've helps you optimize AI agent performance simultaneously across multiple dimensions, including quality, latency, cost, and safety. We've equips developers with a playground for prototyping and testing agent and application LLM calls. Interactive trace trees, like we saw a moment ago, that organize and display all of the most important information to investigate your agent at work. a flexible framework for conducting the rigorous evaluation process required for building bulletproof enterprisegrade AI agents and finally guardrails to protect your brand, your agent and your end users. Weave is purpose-built to support every step of the agent development workflow. Examining our financial research agent provided us with a glimpse into the power of weave when it comes to the iterate and deploy steps. Let's rewind things for a moment and take a look at the prototype step. Building the first version of our AI agent. The process of developing an agent has got to start somewhere. So, let's start at the first step in our financial research agent. Generating search queries to conduct the research required to write a report on a requested financial topic. We'll start in a notebook with some Python code. This is a single task in our agentic workflow. An LLM call requesting those queries that we can then assign to individual financial search agents to perform searches and retrieve relevant information. As you can see, we've added our one line of code initializing weave and also a weave.op decorator to enhance our collected trace data. And after we've executed our call, we can jump directly over to weave and see the call here on the traces page. Clicking on generate searches, we can see a list of all the queries that have been generated. And we can also click on the LLM call itself and then move that into the playground. for further investigation and quick iteration. And when it's time to move on from prototyping into considering which LLM and agent candidates might belong in production, we use weave evaluations. Before we deploy an agent into production, we'll run evaluations on the entire endtoend workflow. But first, we need to start with evaluating each task individually to make sure each part performs as required and contributes successfully to the whole. Here in our notebook, you can see that we're calling three separate scorers to evaluate our generate searches task. Search value score, which evaluates the value and topical relevance of each query using an LLM as a judge. search recency score which evaluates the temporal relevance also using an LLM as a judge as recency and time sensitivity are almost always important factors in financial research and finally search count which is really just a count of how many queries are generated here we're going to evaluate our agent using three separate models used for our generate searches call GPT4 mini GBT40 and 03 mini and once we're done running the scripts, we can jump back over to Weave to check out the results. This is the Weave evaluations page. Here we can see all of the important data from our evaluations runs. But what really brings this information to life is the ability to compare evaluations. Our comparison here shows us that 03 Mini performs the best over all of our evaluation metrics. As production candidates increase, this type of interface that allows you to compare and diff every single option is yet another example of how we've organizes and puts all of the data that you need to iterate, optimize, and make critical deployment decisions right at your fingertips. In addition to the evaluation process, our agentic workflow also includes a verification agent that acts as a meticulous auditor called upon to verify that the report is internally consistent, clearly sourced, and makes no unsupported claims. The presence of the verification agent provides an added layer of protection against delivering results that may be problematic into the hands of our analysts. In the case of the report we requested at the beginning of the demo, the verification agent has given it a passing score while also pointing out that there's still room for improvement. In situations where the verification agent gives a failing score and determines that a report is inadequate, we can add to our agents planner instructions to discard the report and begin a new or direct our financial analyst to ask an administrator for assistance. These verification pass fail scores can also be easily recorded inside of weave and used to filter the data for further analysis or added to future data sets for evaluation or training purposes. While our demo today is focused on W&B weave and the prototype, iterate and deploy steps. The weights and biases AI developer platform plays an instrumental role at every point along the route while circumnavigating the agent development workflow. And as we continue to observe our agents in production, we're always thinking about potential updates to improve the application and the end-user experience. We refine, optimize, and add new features. And using Weave, it'll be easy to measure the impact and performance of these enhancements. Thanks so much for your time, and we invite you to sign up for Weave today, the first step in building and deploying your AI agents and applications with confidence. [Music] [Music]
Original Description
Discover how W&B Weave helps developers build, evaluate, and monitor autonomous AI agents. In this video demonstration, we'll show you how Weave provides the necessary tools for rapidly iterating and optimizing agents so they can be deployed into production with confidence.
*Learn more* →https://wandb.ai/site/agents
*Subscribe to Weights & Biases* → https://bit.ly/45BCkYz
⏳Timestamps:
0:00 Defining AI agents
1:20 An AI agent in action
2:22 Examining an agentic workflow
3:15 AI agent observability with Weave
5:37 Productionizing AI agents using Weave
7:29 The agent development workflow
7:54 Prototyping AI agents
9:22 Evaluating AI agents
10:48 Comparing AI agent evaluation results using Weave
11:15 The benefits of built-in verification agents
12:21 Conclusion and invitation to try W&B Weave
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Weights & Biases · Weights & Biases · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
0. What is machine learning?
Weights & Biases
1. Build Your First Machine Learning Model
Weights & Biases
Intro to ML: Course Overview
Weights & Biases
2. Multi-Layer Perceptrons
Weights & Biases
3. Convolutional Neural Networks
Weights & Biases
Weights & Biases at OpenAI
Weights & Biases
Why Experiment Tracking is Crucial to OpenAI
Weights & Biases
4. Autoencoders
Weights & Biases
5. Sentiment Analysis
Weights & Biases
6. Recurrent Neural Networks [RNNs]
Weights & Biases
7. Text Generation using LSTMs and GRUs
Weights & Biases
8. Text Classification Using Convolutional Neural Networks
Weights & Biases
9. Hybrid LSTMs [Long Short-Term Memory]
Weights & Biases
Toyota Research Institute on Experiment Tracking with Weights & Biases
Weights & Biases
Weights and Biases - Developer Tools for Deep Learning
Weights & Biases
Introducing Weights & Biases
Weights & Biases
10. Seq2Seq Models
Weights & Biases
11. Transfer Learning for Domain-Specific Image Classification with Small Datasets
Weights & Biases
12. One-shot learning for teaching neural networks to classify objects never seen before
Weights & Biases
13. Speech Recognition with Convolutional Neural Networks in Keras/TensorFlow
Weights & Biases
14. Data Augmentation | Keras
Weights & Biases
15. Batch Size and Learning Rate in CNNs
Weights & Biases
Applied Deep Learning Fellowship Overview and Project Selection with Josh Tobin (2019)
Weights & Biases
Grading Rubric for AI Applications with Sergey Karayev (2019)
Weights & Biases
16. Video Frame Prediction using CNNs and LSTMs (2019)
Weights & Biases
Image to LaTeX - Applied Deep Learning Fellowship (2019)
Weights & Biases
17. Build and Deploy an Emotion Classifier (2019)
Weights & Biases
Applied Deep Learning - Data Management with Josh Tobin (2019)
Weights & Biases
Snorkel: Programming Training Data with Paroma Varma of Stanford University (2019)
Weights & Biases
Applied Deep Learning - Troubleshooting and Debugging with Josh Tobin (2019)
Weights & Biases
Troubleshooting and Iterating ML Models with Lee Redden (2019)
Weights & Biases
Designing a Machine Learning Project with Neal Khosla (2019)
Weights & Biases
Lukas Beiwald on ML Tools and Experiment Management (2019)
Weights & Biases
Building Machine Learning Teams with Josh Tobin (2019)
Weights & Biases
Pieter Abeel on Potential Deep Learning Research Directions (2019)
Weights & Biases
Testing and Deployment of Deep Learning Models with Josh Tobin (2019)
Weights & Biases
Five Lessons for Team-Oriented Research with Peter Welder (2019)
Weights & Biases
Applied Deep Learning - Rosanne Liu on AI Research (2019)
Weights & Biases
Making the Mid-career Leap from Urban Design to Deep Learning/Data Science
Weights & Biases
Organizing ML projects — W&B walkthrough (2020)
Weights & Biases
Brandon Rohrer — Machine Learning in Production for Robots
Weights & Biases
Nicolas Koumchatzky — Machine Learning in Production for Self-Driving Cars
Weights & Biases
My experiments with Reinforcement Learning with Jariullah Safi
Weights & Biases
Applications of Machine Learning to COVID-19 Research with Isaac Godfried
Weights & Biases
Testing Machine Learning Models with Eric Schles
Weights & Biases
How Linear Algebra is not like Algebra with Charles Frye
Weights & Biases
Predicting Protein Structures using Deep Learning with Jonathan King
Weights & Biases
Rachael Tatman — Conversational AI and Linguistics
Weights & Biases
Reformer by Han Lee
Weights & Biases
Sequence Models with Pujaa Rajan
Weights & Biases
GitHub Actions & Machine Learning Workflows with Hamel Husain
Weights & Biases
Look Mom, No Indices! Vector Calculus with the Fréchet Derivative by Charles Frye
Weights & Biases
Jack Clark — Building Trustworthy AI Systems
Weights & Biases
Surprising Utility of Surprise: Why ML Uses Negative Log Probabilities - Charles Frye
Weights & Biases
Track your machine learning experiments locally, with W&B Local - Chris Van Pelt
Weights & Biases
Antipatterns in open source research code with Jariullah Safi
Weights & Biases
Attention for time series forecasting & COVID predictions - Isaac Godfried
Weights & Biases
Made with ML - Goku Mohandas
Weights & Biases
Angela & Danielle — Designing ML Models for Millions of Consumer Robots
Weights & Biases
Deep Learning Salon by Weights & Biases
Weights & Biases
More on: Agent Foundations
View skill →Related Reads
📰
📰
📰
📰
How MADDPG Combines Deep Learning with Multi-Agent Strategies
Medium · Deep Learning
The one-page contract I make every autonomous AI agent sign (template inside)
Dev.to AI
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
ArXiv cs.AI
Deterministic Replay for AI Agent Systems
ArXiv cs.AI
Chapters (11)
Defining AI agents
1:20
An AI agent in action
2:22
Examining an agentic workflow
3:15
AI agent observability with Weave
5:37
Productionizing AI agents using Weave
7:29
The agent development workflow
7:54
Prototyping AI agents
9:22
Evaluating AI agents
10:48
Comparing AI agent evaluation results using Weave
11:15
The benefits of built-in verification agents
12:21
Conclusion and invitation to try W&B Weave
🎓
Tutor Explanation
DeepCamp AI