Maximize Your Productivity with LLMs: Task Utility Explained
Key Takeaways
The video discusses the concept of task utility in Large Language Models (LLMs) and how it can be used to evaluate and optimize LLM-powered applications, with a focus on agent-based systems and retrieval augmented generation, using tools such as Autogen, Aogen U, and Agent Eval.
Full Transcript
we're bringing on Julia here from mulon Julia has a background working on uh researching user-based scale satisfa satisfaction metrics and methods for the last 15 years dating to uh web search scenarios and spent six years conducting research on uh agent related Concepts in Microsoft and now works building uh a at mulon uh who has a tagline of building agents that complete tasks start to end so I'm very curious to see what uh Julia has to share with us with that I'll let you take it away and I'll come back in for Q&A at the end thank you so much for the introduction um happy to be here and uh share some of my uh work on uh ENT the relations so and um just a highlight uh what you will get out of this uh talk today so uh if you're interested to build a l owered applications was basically like uh in a gentic space and uh you building for where is the means so you want to understand what potential uh metrics you need to optimize your application for and then it's actually uh a lot of work to figure out uh what needs to be done for example like if you look at an examp a great example of such an application such a web search has been a lot of work and research uh done uh by a huge Community to even figure out what's kind of a right metric uh to optimize for from the user perspectives because in the end of the day we all building those type of things for the users and um I was lucky to be part of uh autogen so it's a probably if you uh looked in a gentic space probably heard about this open source programming framework um which is has been released uh in uh September of the last year and I think it was one of the best and attract a lot of attention so and uh I was just blown away by the idea how many people are trying to use the agents um L empowered agents to solve all kind of tasks all kind of tasks you could see like especially like on Discord Community we could see like people were using it to build like teachers to build IST and uh things like that but the uh difficult part uh in machine learning is actually figuring out what's your objectives what basically you uh kind of build in what you want to optimize for and what you uh uh want to design as a met as a metric so it's uh relatively easy if you think about such task like for example as uh object recognitions because you can just get the um ground truths uh and label that and just focus on making sure that your algorithm is working well but if you're building something for the end user you need to care about how your end user say satisfied or not satisfied with an experience and what your application is even uh done for and um so kind of the motivation to get into this idea of a kind of uh scalable approach for evaluating um edent performance or like not even evaluating but even understanding what kind of utility it may or may not bring to the users so uh we decided to uh start to explore this because we have seen how many uh different types of um uh developers all over the world that building the applications for the users and they uh have to have a tools to figure out if the applications uh if the agent are building in actually achieving the uh uh achieving the needs user actually uh kind of highlighted and um so again most of the time uh these type of things are um done in like in machine learning settings when you have a uh sort of ground truths you uh collected you labeled but it's not always possible especially with such a rapid way of technology and even like again uh for that you need to at least understand what's your objectives so and um there's also a lot of research has been done already to show that that llm itself can be a efficient alternative to uh get a idea about what your users actually want and help you to evaluate the um uh systems you build in and you need to do it fast and it's a relatively cost-efficient way of doing this so uh I will not go deep into this because it's not part of the stock but if you're interested like you can write a lot of papers where you could maybe highlight this llm as a judge so and this uh process needs to be automated uh and especially it's important for the again for the tasks when you sometimes do not have don't have any ground truths and you want to figure out uh this by uh going through your data so and we basically another approach another important aspect why we really want to bring agent devel there is like also empowering the developers and make them aware uh What uh potential utility uh of the applications uh they designing uh can do to the end users and make them um aware about that and hopefully care so that they while they will be improving and changing their uh agent system behind it so they actually know they are not acting in the dark but they have a tool to figure out if it's uh getting better or worse so if you let me be a bit more structured on this uh aspect it's like um so we have uh seen a lot of um tasks uh which uh is users are trying to solve using the llm or LM powered applications and there are tasks where success is not clearly defined because success is usually for for us as machine learning researchers it's something like uh great if it's on binary scale so if we have completed the task or not completed the task so but for example if you see the uh assistant is writing email for you and then you can of copy this email and change it here and there so like what exactly was the success here so uh that's uh still own like um interesting direction for the research but there are tasks where success are clearly defined for example of the goal of your like LM part application is to solve the mass problem so like one of the um one of those benchmarks we all care about so like the uh success is clearly defined so because we have a ground truth we know the right answer we can judge if uh the uh solution was correct or not so another aspect of those test is that when you do have a a kind of preferred optimal solution uh for this and there are some cases where there's multiple Solutions which lead to the same uh sometimes incorrect answer and you would would need to figure it out which one is actually better uh for you or not so why I'm bringing up uh that here because uh what we going to uh discuss next is a new way of uh defining what's uh utility uh the LM powerered application can bring to the end users and we try to be uh as domain agnostic as possible so meaning we can provideed with a concrete Benchmark and uh so the whole idea about benchmarks at the moment is Pro also needs to be discussed in separate talk so but uh we while we building this approach we want to kind of show that it makes uh uh sense so we need to kind of start somewhere so and we decided so obviously that's going to be much easier to start with a tasks where success is clearly defined and see what we can do uh for this type of tasks um uh uh in this space and uh so uh so for this particular presentation uh I will be using uh the task of solving the mathematical problems as one of those so which is clearly defined and um uh so like we tried agent well on various types of tasks even not where the success was defined afterwards so like you can look at up this uh in the paper or online so first uh we need to uh reconsider uh the idea that you need to define the utility because you're not building the application just for the application you want to help users to solve some particular tasks so and um usually uh so you have those uh criterias in mind like that's uh you building the solution which is doing something faster than another solution or it's kind of L some more effort so uh but there might be something uh which you cannot really uh even um uh look at at the moment so and usually when you do that for the new tasks uh like for example how to evalate the uh teacher performance you have to define a lot of criterias and usually takes uh like used to take a lot of effort and a lot of ongoing user research uh uh and so on and so forth so what we go here is that like most of the developers do not have this type of um um support time or even uh kind of proper education to do those user research so can we uh Define the critic Ed which will help uh uh the developers with uh defining the criterias uh which the application is uh uh needs to be kind of evaluated and that's basically those criterias express some kind of utility to the end user and uh the quantifier agents can assess uh how well your particular solution is doing according to defined criterias so we basically trying to Define kind of multi-dimensional task utility and just to give you uh like thing out of uh scope so I picked a math example for the uh for the reasons that I think we all probably a bit um the main uh expert in MTH here because um we all studied it at school or maybe at some other levels so basically here you see the uh uh criterias for the MTH problem which was justified by the critic agent so that's basically the uh agent uh which you provide was um uh prompt uh where you ask agent to uh kind of explain what the task is and you provide one successful example and one unsuccessful example of your uh agent solving a particular problem and so here I just picked like uh four criterias and it's a descript description and accepted values so and it's all uh the output of the uh can uh the agent so nothing of that I came up uh myself so I can just at this point it's like especially if like for example I designed a new um a new application for a new domain and I have seen for example we tried the experiment for the robotics space uh so where I I was I'm not an expert so and actually like the critic agent helped us to figure out like quite a complete set of criterias which was verified by the domain expert later on so like like uh that's basically saves us a lot of time on figuring out what are those important criterias which uh are very important for this domain so and again like you can do it very easily with just a Critic agent which address you this table it's very important to follow the accepted values as suggested by the critic agents here so like all this zero ones and two that's what we did to kind of uh clarified and like use it for the later uh histograms but uh in general so like better uh stick with the language of it so just uh if you have uh I can go on longer ex explanations why but just like believe me that was the best way to to go to make sure that like you provide for the quantify edge of the um uh textual description of this so basically uh that's all uh comes for free and by the way so you would say like how is it important that was really important for example one of the first criteria uh which was shown uh for the mass problems was um the uh code quality and I was like what kind of code quality and then I realized that aogen U so like one of the uh baselines which we use produce a a Pon con to solve this problems that's why we're kind of achieving very high quality so very high uh uh so like basically it's uh doing the um uh tasks right uh the question is if you want to actually solve it uh you then provide like writing the code so for example if you're trying to do something like a teacher agent is it that something you want to do or not again it's up to you it's a uh your application so we're just here to help you evaluate and this Define what the task utility can be so and um another thing you might ask me is that like the quantify agent in this case uh sees unseen um examples of uh your uh agent and uh tries to quantify them so basically like if you define the new solution or provide a new Baseline and you stick with kind of task utility you want to be uh following this type of criteria so uh the quantify agents can uh help you to uh assess the Unseen examples and see if like you improving on some of those criterias so and desperately needed here at the verifier agent because uh the verifier agent can tell you which criterias are good or maybe like it's better to say which criterias out of what you have is like uh can be easily assessed by the quantify agent or not so and we specifi the uh verifier agent which uh will help help you to filter the criterias which are uh actually uh stable and robust and also they are very so they supposed to be um uh good with adversarial examples and I will show you a bit of uh example uh later so you can down like here I think you can download the paper and by the way the AL all is part of the alen now you just go run it so but so that's the main part of the talk you you uh may want to remember so imagine you are designing your uh LM powerered application you use agents which are communicating to each other to solve the task they produce a lot of logs for you so and you can basically based on this logs figure it out what kind of criteria are important for your task uh and hopefully for your end users as well you can use a uh uh quantifier agents to uh kind of assess the quality of each of unseen uh or unseen uh Solutions and then us then the verifier agents you can just uh filter the criterias which we think are not really St that's possible because uh there's a lot of things needs to be done still to figure out how uh well this quantifier agent works and you can do it for any type of domain and you don't have to uh get the um uh Crown TRS uh uh for this uh so if you have that's great U but if you don't then just want to explore it there something that you can in couple of minutes set up and uh get the first idea what's your uh users can uh be using it for so here just give you also a bit of an example and let's come back to this idea of like we we developing something new this is a new approach of how you can uh do the sort of assessment we call it assessment because it's a bit over um stretch called evaluation uh in my opinion so but like that's why we split we start with a task where we can split the successful and unsuccessful uh tasks and so here what you see you see the same criteria so it's only four usually it's up to 25 criterias you can get uh for majority of tasks and uh uh you could see like three baselines it's like react uh uh vanilla sorus gp4 and aogen and you could see that for example for clarity efficiency and completeness you have also in line that successful tasks actually higher on averaged uh for each criteria that's kind of the things we were trying to make sure that we following so just because again we're doing something new we have to have some kind of uh ways of uh verifying if we're doing it right or not so uh but you can see that there's a variation in terms of the how the uh various types of um uh criterias are kind of assessed by uh different uh um baselines and that's kind of can be like then up to you to decide if you want to go with what one of them but also important part is like to keep those assessments and if you change your uh agents underlying agents you can uh reiterate the process and see first of all there's a possibility that your new uh Baseline can bring new criteria into the place because uh kind of few solution sometimes brings some of these criterias and um um you could see if that's something you you want to go with however this uh error analysis doesn't look really good for us so like that's why we think that quantifier couldn't do a job uh to assess uh the proper uh uh kind of the proper way the error analysis uh and that's what the verifier is for because a verifier will tell you that maybe aor analysis is a good criteria you want to have but we just cannot assess it in the current state so and that's definitely more research needs to be done uh towards figuring out if uh can do something about it so and um this is uh another important part to uh figure it out is that there a task based uh criteria and there solution based criteria so uh sometimes if you just describe to L what you want to evaluate based on a task it can suggest you different criterias but providing the successful and unsuccessful example can uh open up the types of again like pre if you just described to LM like what's a IAL for solving the mass problem the code efficiency would not be part of it at least like at the times when we tried it but since uh now sometimes we use the code to solve the mass problems that's uh kind of that's part of it now so but uh it's also kind of important what you could see here that uh you uh at some point kind of uh that would be suspicious if our critic will uh like uh uh give us the more um kind of uh iterations we run the more criteria is going to suggest so like it looks like at some point we can PL one a number of criterias which is a good sign so like basically there's a a limited number of deaths and from my own my own intake here is like I would not worry too much about uh the accuracy of this criteria sometimes people say like are they going to be correct it's only uh like in the worst case 25 criteria 30 criterias per domain as a domain expert you can easily verify some of those and you can remove which the one which you're not interested I would be more concerned figuring out how complete is this Set uh so and that's research needs to be done on this uh in this area as well but another intake here please uh try to as soon as you propose a new solution around your critic you would be surprised because the critic might discover new criterias which is introduced by your uh solution so and this is another uh kind of uh ideas about how can you uh see uh yes again we have this uh um um Advantage here that we had like successful and unsuccessful cases so and then just to see um the distribution of a quantifier output uh for uh successful you can see there a dark blue and failed cases and for example again for a analysis you could see that uh basically the scores are are kind of almost uh the same uh on average that's not what you want to see so uh but again we learned it based on uh the examples how to uh for Bas on examples where we have success criterias and like using the very far you can filter the criterias which are unstable even if you don't have uh don't know the uh success of it so and um another thing is that like was interesting to uh do with figuring out if the uh sort of worse examples or Worse solution is actually uh show to be wors and for different criterias it's not such an easy task because uh again we just have here like either successful or not successful Solutions that's why uh so like you didn't have a a gradation so like as uh you could see here this average Welles which is corresponded to the um um the categories I showed you in the beginning so what we did we just uh take and introduce the noise in an existing solution and then uh the hypothesis was that the samples which would be uh having contain in noise so disturb samples as we could see here should get lower score or L and quantifier uh for in comparison to the same um examples uh but without noise and so we could see that this is uh actually the case so that means that you can sort of trust the idea that quantifier will uh rank higher the examples uh where like which are like sort of of a better quality uh for for uh various types of criteria again this is important because we're uh trying to set up something absolutely new uh here and uh as a conclusions uh so we introduced a novel uh framework we call it um agent eval uh so it's a desire to quickly Val evalate LM part agentic application here you have a QR code uh you can go to the blog post with uh on um as uh part of the aogen library and just try agent evolve for yourl and power duplication also uh actually that's uh that's based on academic paper which will be presented tomorrow I think at emop as well so and um we believe that this can be used for any type of uh agentic applications where you have logs which you can analyze and see how your agents is uh behaving and that should also should be very scalable and it's cost efficient as well so you don't really need to uh run um even do a lot of llm calls to get your first sense of various types of uh criterias and um um it can give you kind of interesting outline is like what that potential utility can be so for example for this mathematical uh math application some of the criterias was both the solution was ver bothos and maybe you want verbal Solutions especially if you designing something like a teacher so and the beauty of this is that like then having this in place you can actually keep optimizing your um agents for providing the more ver both Solutions or you can even give it to uh Outsource it to the users and ask them what kind of uh solutions they want to have what kind of utility they have but having those criterias and having a good way of figuring out if it's actually uh Your solution is following those criterias using the quantifier can help you to uh kind of optimize your uh application towards the user needs so uh and another thing is that like uh as I as I said like agent Devol helps you to uncover your capabilities and uh seeing how they evolved over time because if you propose a new solution there's a possibility that uh that will bring a new criteria for you as well so keep an eye on that and we hope that's going to help uh the developers all um um developer Community to figure out how to uh assess this solution and align them with the needs or with the criteria they want to uh kind of uh uh associate the solutions uh with so and this uh definitely uh potentials a lot of potentials work this one one of the first directions uh which we can go and um I'm just uh super happy with here today to share this work because I think that's potentially it's uh can un cover um the ways for developers to actually figure out their strongs uh strengths and weaknesses of the agentic applications based on real data based on the um inter uh kind of interactions and uh just also bring awareness that we need to optimize our applications not only towards the obvious uh metrics like um latency uh the even success but also this underlying uh multi-dimensional way of how users receive our application and have a way to uh influence that through this and thank you so much for your attention that's I pretty much completing uh my talk today awesome thank you so much much for coming uh fortunately we don't have time for questions uh so you know uh at this time uh everyone can Mosey on over back to track one to view the closing but uh really appreciate everyone coming and viewing and really appreciate your time Julia uh so yeah thank thank you so much if you have any questions uh please reach out to me are um in LinkedIn DM on Twitter is Al open so I'm happy to follow up on any questions regarding this work uh and again it's available as a part of the of the join library and the more extensive explanation as a part of the um paper yeah awesome thank you so much take care thank you
Original Description
//Abstract
The rapid development of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents, assisting humans in their daily tasks. However, a significant gap remains in assessing to what extent LLM-powered applications genuinely enhance user experience and task execution efficiency.
This highlights the need to verify utility of LLM-powered applications, particularly by ensuring alignment between the application's functionality and end-user needs. We introduce AgentEval, a novel framework designed to simplify the utility verification process by automatically proposing a set of criteria tailored to the unique purpose of any given application. This allows for a comprehensive assessment, quantifying the utility of an application against the suggested criteria.
//Bio
Dr. Julia Kiseleva joined MultiOn to advance their research in building safe and reliable AI agents, with special attention to delivering high-quality solutions. Previously, Julia led projects seeking to answer important questions, such as how to build and evaluate interactive agents with humans in the loop. Notably, she contributed to the NeurIPS competition on Interactive Grounded Language Understanding (IGLU) and developed new strategies for evaluating agentic systems (AgentEval) and many other initiatives, always driving for user-driven yet scalable evaluation of interactive systems.
A Prosus | MLOps Community Production
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from MLOps.community · MLOps.community · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Our 1st MLOps Meetup // Luke Marsden // MLOps Meetup #1
MLOps.community
Remote Collaboration as a Data Scientist
MLOps.community
MLOps Manifesto with Luke Marsden from Dotscience
MLOps.community
MLOps lifecycle description
MLOps.community
What Does Best in Class AI/ML Governance Look Like in Fin Services? // Charles Radclyffe // MLOps #2
MLOps.community
Life purpose and too many spreadsheets
MLOps.community
Explainability, Black boxes and EU white paper on reproducibility
MLOps.community
Hierarchy of Machine Learning Needs // Phil Winder // MLOps Meetup #3
MLOps.community
Automatically Retrain Machine Learning Models? Are best practices worth it?
MLOps.community
Building an MLOps Team? Key ideas to keep in mind
MLOps.community
Hierarchy of MLOps Needs
MLOps.community
Bare necessities for getting an ML model into production
MLOps.community
MLOps and Monitoring
MLOps.community
How Phil Winder got into Data Science and Software Engineering
MLOps.community
Provenance and Reproducibility in Machine Learning; what is it and why you need it?
MLOps.community
Friction Between Data Scientists and Software Engineers
MLOps.community
MLOps Problems in different size companies
MLOps.community
ML tooling in large companies
MLOps.community
ML Platforms - The build vs buy question
MLOps.community
ML Services Gateway at SurveyMonkey
MLOps.community
Message buses, Async and sync architecture
MLOps.community
MLOps #4: Shubhi Jain - Building an ML Platform @SurveyMonkey
MLOps.community
Hybrid Data Science Teams @SurveyMonkey
MLOps.community
How do you handle ML version control at SurveyMonkey
MLOps.community
Doing ML with Personal Information
MLOps.community
Evolution of the ML feature store @SurveyMonkey
MLOps.community
Developing a Machine Learning Feature Store
MLOps.community
Auto retrain ML models is not the question
MLOps.community
3 key parts to Machine Learning monitoring
MLOps.community
MLOps Meetup #6: Mid-Scale Production Feature Engineering with Dr. Venkata Pingali
MLOps.community
MLOps meetup #5 High Stakes ML: Active Failures, Latent Factors with Flavio Clesio
MLOps.community
MLOps: Airflow Pros and Cons
MLOps.community
Specific challenges in Machine Learning
MLOps.community
Current State Of Machine Learning
MLOps.community
Humans in the Loop are a defining factor in Machine Learning
MLOps.community
Learning from real life Machine Learning failures
MLOps.community
Survivorship Bias in machine learning tutorials
MLOps.community
Swiss Cheese model in Machine Learning
MLOps.community
Resume driven development in Machine learning & software engineering
MLOps.community
Who has the highest standards in ML?
MLOps.community
Venkata Pingali of Scribble Data Thoughts on the Current State of Machine Learning
MLOps.community
Dependable data and being able to Trust in your Data with Venkata Pengali of Scribble Data
MLOps.community
Speed, Trust, Evolution and Scale in MLOps
MLOps.community
More difficult transition for data scientists to become ML engineers
MLOps.community
How many models in prod til I need a dedicated ML platform?
MLOps.community
Deeper thinking from data scientists around platform blackholes
MLOps.community
Checkpointing, metadata, and confidence in your data
MLOps.community
Adjacent usecases and multistep feature engineering
MLOps.community
Standardization of Machine Learning tools like in Software Engineering with Venkata Pingali
MLOps.community
Reproducability flaws in end to end Machine Learning debugging
MLOps.community
3rd wave of data scientists
MLOps.community
MLOps meetup #7 Alex Spanos // TrueLayer 's MLOps Pipeline
MLOps.community
MLOps Meetup #8 Optimizing Your ML Workflow with Kubeflow 1.0
MLOps.community
Are Kubeflow and Airflow complementary?
MLOps.community
Why Kubeflow gained so much traction=open community
MLOps.community
Who decides the dirrection of Kubeflow
MLOps.community
What do Kubeflow and Arrikto do and how do they work together?
MLOps.community
Versioning your ML steps with Kubeflow
MLOps.community
Machine Learning Lifecycles//Perception vs Reality
MLOps.community
Kubeflow vs SageMaker in Machine Learning
MLOps.community
More on: LLM Foundations
View skill →Related Reads
📰
📰
📰
📰
Will Developers Need LLM Integration Skills in 2026 for Success?
Dev.to AI
I Trained a 471M-Parameter Language Model From Scratch on One RTX 4090 in 100 Hours.
Medium · LLM
Masking PII Without Losing It
Medium · LLM
Build a Career in Artificial Intelligence : AI Mastery Course in Telugu
Dev.to AI
🎓
Tutor Explanation
DeepCamp AI