Behind the Scenes: Accelerating the AI Agent DevOps Lifecycle with End-to-End | LIVE159

Microsoft Developer · Beginner ·🤖 AI Agents & Automation ·1mo ago

Key Takeaways

Accelerates AI agent DevOps lifecycle with end-to-end observability and optimization

Full Transcript

[music] >> Go. All right. Thank you all very much for coming to the session and I'm super excited to be here with two of my favorite people. So, uh Vivek and Felicia, do you want to just give me a quick introduction to who you are and what you do? >> Hey Nidhi, yeah, super excited to be here. I'm Felicia. I work on the Foundry Observability team. >> Hey, I'm Vivek and I also work on Foundry. >> And I know we're talking a lot about agent optimization and like why it's so important. So, maybe you can kick this off for us Vivek by kind of telling us what makes it so difficult for us to optimize an agent today. >> Yeah, that's a good one and uh agents are different from regular piece of software. Uh the properties that makes agents so useful, like they're stateful, they they're long horizon, they can plan, they can course correct, they interact with their environment through tools, are also the ones that makes it very hard to test. >> Mhm. >> so, regular piece of software, I generally have to find the right piece of logs and the right piece of software and I can just trace through what's going on and I can understand. >> Uh >> With agents, that's no longer true because as a developer, I'm not coding everything as is. That's the feature of agents. They also interact with the environment, so their state the environment state and agent state drives its behavior. So, it makes it incredibly hard to test, to reproduce. Uh and those are all the problems where we are trying to solve with Foundry agent platform, too. All the way from build, run, and operate, uh where you can trace, eval, and optimize all in one platform. >> I think actually and we've been doing this for day now and like talking to a bunch of developers and having like a unified platform with end-to-end observability all the way from your planning to production is amazing. But, he mentioned logs. So, how exactly do we go from like looking at these logs to getting to the point where we >> Yeah, I think observability or like actually testing agents is really hard. I think that's that's sort of where Vivek is going and that's where like Foundry really shines in some ways is like being able to understand how to test something that's non-deterministic is like the whole new science behind evaluations, right? Evaluators essentially are these you know, they they they rely on queries between like you like sample queries or like sample data and you know, they follow a pattern that they that you then use a non-deterministic LLM behind the scenes to then ask, "Hey, my agent, how does it match up against something that I think should be the kind of answer?" So, from a from a very base perspective, evaluators are sort of your new test-driven development is what I like to call it, right? And um we can use them both in the inner loop for development, you know, where you evaluate on top of data sets and you can use them in the outer loop where you evaluate on top of like live production traces coming in. And that connection helps you keep up with the health of your agent. >> Yeah, but I I was also going to say that one of the things that really struck out was the email rubrics. >> Yes. Oh, yeah. People are super excited about that. >> Yes, I do for us. >> The email rubrics is like a new type of evaluator that Foundry has come up with and gives you this multi-dimensional way to evaluate, which is specific to your agent and it gives you like these weighted evaluators and they tackle a lot of different dimensions and you'll see that the coverage is so good that no matter what type of you know, thing metric you're trying to track, it is able to handle that pretty well and you know, guardrail against some edge cases, yeah. >> I will tell you, people were fascinated by email rubrics and I can't wait to try them out. But I think we kind of talked about why agent devops is changing and we've talked about why evaluation-driven development is important for developers. But there's this one left last thing, right? We all know that the first time we build an agent, we can make it work. But once it's in production, it needs to keep changing. So, what are the things that we need to think about once it's in production to keep continuously optimizing it. Maybe I'll start with the inner loop. >> Yeah, I mean I I think I I I heat this up a little bit, but I think the thing that you want to make sure is that again, you are testing against any regressions, right? So, whenever you make any It's just the same principles that we use for regular apps. It's just you are now applying the non-determinism layer. So, the things that you are testing are just slightly different. So, you're doing the same thing and the way to maintain the piety of that is to maintain the health of your data sets. There's this concept like golden data sets and these concepts that come in that you test anything you do against certain sets of evaluators that Foundry will recommend, by the way, to you. And then you test against that and then you hand it to production and then >> Yeah, and I will say that I completely blame Felicia for this, but she made me kind of learn about Foundry skills. They were amazing. But that brings us to the highlight of this like when in production What do we do? >> Yeah, absolutely. So, the vision is that the more agent runs, the better it gets. >> Yeah. >> And to enable enable it, it really starts with inner loop where you you define what good looks like through evaluations, rubrics, and then the production traces sort of drive your representative set of tasks. And if you can use those tasks and then rubrics to then reflect and see that what should change in this agent and then assess if that change is actually working based on again the data sets that created through the traces and the rubrics that are provided, then you can continuously improve that agent. That's where the Foundry agent optimizer is working. And the the goal really is to have this sort of this continuously improving agents. The more sessions they see, the better they become. >> Yeah, and I I have to give a shout out. If you haven't seen it, you should definitely go check out their breakout. It was BRK252. Ask me how I know. But it was the most awesome demo that you did. An agent optimizer I think is going to be one of the things everyone should try out. Is it in private preview public? >> It's in gated preview and it'll be coming public preview by end of this month, but you can still sign up and you can check out. >> So remember his name send him all your feedback cuz he's going to personally fix it all. >> [laughter] >> But I think that those are the main things that I heard in the end-to-end observability story. So agent optimizer, foundry skills, evals and I think the other thing that I heard is we can now do evaluation of both single turn and multi-turn conversations. >> Yes, you can give the entire conversational context, right? Like so you multi-turn and then you also have the user simulations that you can check on top of your traces to see exactly if there is a trace that's run against your eval and there's something that's gone wrong. You can like with a quick click actually find the representative query of what that was the you know harmful trace or whatever. >> Yeah, check out trace replays. You will not regret it. The coolest thing ever. >> I love how excited you are by this Nitesh. [laughter] So let me ask you have What is the best agent you have made so far? >> Oh, you're just setting me up. Okay, so because she said so you got to go check out the lab. It's not the best agent. The best agent is the agent you learn from, right? Like you got to have an agent that breaks so you can fix it and you learn by fixing it. So that's my favorite agent. What did you expect me to say James Bond or something? That won't [laughter] work. But we have about I think a minute left. So I'm going to put you on the spot. You're not allowed to choose your own products because that's what they wanted to do. What is the one thing you got out of Build? Like what was one thing that stood out for you at Build this year? >> Yeah, and this is my first Build. So I really just like interacting with the developers over here. We've been building for a while. A lot of these pieces of software and thinking about how end-to-end observability and optimization should be, but this talking to the developers one-on-one talking to them like how they are solving their problems and how this fits in is so exciting. That has been the best part. There are so many takeaways that I have, so many LinkedIn connects I have from all of that. Yeah. >> I was going to say we we late named it Build and it's all about builders. How about you, Felicia? >> I really like the food that they're serving. >> [laughter] >> So, I really like how people go around on trays and give you food. I genuinely think it's really fun. I like the venue. It's kind of fun that it's indoor outdoor. Um I feel like a lot of people are seeing the day for the first time in their life. You know, it's kind of just bringing the outdoors to developers. But I I I actually the thing that I actually really like is um um you know, collaborating in person. I think the um it always energizes me to be around people and you know, having a good time with my coworkers backstage. It makes me realize that I'm working with real humans to build agents. >> There you have it, folks. Um thank you all very much and don't forget to go check out the BRK252 uh breakout and we'll see you all. Have a wonderful Build. >> Thank you.

Original Description

In this interview we unpack what it really took to deliver our end‑to‑end “observe → evaluate → optimize” flow covering the end-to-end Agent DevOps lifecycle from inner‑loop offline signals to continuous improvement in production. We’ll share hard-earned lessons on (1) simplifying the getting started experience with out of the box observability powered by context-specific eval rubrics, (2) streamlining the developer experience with guided skill-based flows and (3) leveraging the complete set of inner and outer loop signals for continuous improvement. 𝗦𝗽𝗲𝗮𝗸𝗲𝗿𝘀: * Vivek Bhadauria * Filisha Shah * Nitya Narasimhan 𝗦𝗲𝘀𝘀𝗶𝗼𝗻 𝗜𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗼𝗻: This is one of many sessions from the Microsoft Build 2026 event. View even more sessions on-demand and learn about Microsoft Build at https://build.microsoft.com LIVE159 | English (US) Broadcast Stage #MSBuild Chapters: 0:00 - Why Agents Differ from Regular Software 00:02:44 - Using Evaluations to Monitor Agent Health 00:03:39 - Discussion on Continuous Optimization of Agents in Production 00:05:35 - Mention of BRK252 demo showcasing Agent Optimizer 00:05:49 - Agent Optimizer preview details - gated now, public by month-end 00:06:20 - Using user simulations and Trace Replace to identify harmful traces 00:06:47 - Discussion on favorite AI agent and learning through broken agents 00:07:33 - Conversation about developer collaboration and end-to-end observability 00:07:48 - Host emphasizes Build as an event for builders and asks colleague’s opinion
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
We put adversarial agent testing directly in Claude Code and Cursor
Learn how to test AI agents using adversarial testing directly in Claude Code and Cursor, streamlining your development process
Dev.to · Sofia Aliferi
📰
The Agentic OS: Apple's AI Strategy for Enterprise Platform Teams
Learn how Apple's Agentic OS embodies a local-first AI strategy and its implications for enterprise platform teams
Dev.to · Omnithium
📰
AI Agent Skill Registry: Stop Prompt Sprawl Before Workflows Break
Learn to design an AI agent skill registry to prevent prompt sprawl and workflow breakdowns
Dev.to · Jack M
📰
The Future of Agentic AI in 2026
Learn how multi-agent AI is revolutionizing the field by enabling specialized agents to collaborate and achieve complex tasks
Dev.to · Divyanshi Kulkarni

Chapters (9)

Why Agents Differ from Regular Software
2:44 Using Evaluations to Monitor Agent Health
3:39 Discussion on Continuous Optimization of Agents in Production
5:35 Mention of BRK252 demo showcasing Agent Optimizer
5:49 Agent Optimizer preview details - gated now, public by month-end
6:20 Using user simulations and Trace Replace to identify harmful traces
6:47 Discussion on favorite AI agent and learning through broken agents
7:33 Conversation about developer collaboration and end-to-end observability
7:48 Host emphasizes Build as an event for builders and asks colleague’s opinion
Up next
6 Agentic AI Projects: Every AI Engineer Needs in 2026
Rajeev Kanth | BEPEC
Watch →