Build Specialized AI Agents: Post-GTC Developer Deep Dive
Skills:
Agent Foundations90%Tool Use & Function Calling80%CV Basics70%Modern CV Models70%Generative CV60%
Key Takeaways
The video demonstrates how to build specialized AI agents using NVIDIA's Neotron models, including Neotron Nano2B2VL and NanoVL, and explores the latest advancements in agentic AI for developing intelligent, multimodal systems. It covers various tools and techniques, such as retrieval augmented generation, fine-tuning, and multimodal understanding, and provides insights into responsible AI, guardrail models, and model cards plus.
Full Transcript
What is up everybody? Welcome back to Neatron Labs live stream. Happy Tuesday. Uh let us know in the chat if you saw us at GDC DC. It was great meeting some of you face to face. Uh today we've got uh someone awesome. This is Annie. She's gonna take us straight into our demo. Annie, take us away. Show us what we're here to talk about. Okay, Annie's got a a few things. It looks like her her mic's out of a problem. Yeah, Andy, take us away. Excellent. Hey folks. Um, here I'm excited to uh showcase uh the Neotron Nano2B2VL. It's our second vision uh language model in the Neotron family. And uh this is build.envidia.com. You can uh quickly get started with this model here. I'm going to show you how to interact with it. Um let me try uploading an image right here. And you can go ahead and ask any question. I'm going to turn off reasoning uh because I'm going to ask a simple question. How to describe just describe the image describe the contents in this image. And there you see uh the model um is very much trained on doing amazing um OCR and it's going to give you a very good rich detail of what's inside this image. Uh let's go ahead. Let's go ahead and uh do a little bit more um uh tinkering with this. Here I have a PDF of quarterly earnings of FI 2026 from Nvidia. So I've taken a couple of of these images and trying to prompt the model. So as you saw previously I was able to upload one singular image but this model with its in um more uh improved reasoning capabilities it can do multiple images. So here I'm I'm actually asking a question on uh about six images from the PDF and um I ask here how much did data center business grow in Q2 of FI 2026. Let's go ahead and run the cell. While it's going to give us the answer, let's go ahead and see what the answer is supposed to be. So right here from the data center uh graph, I see that 56% year-over-year growth and 5% quarter overquarter growth. and it gave exactly the right answer. So it pulled the right information among the six images I gave it and gave me the right answer. So one one best practice for you all before you go ahead and test this model um if you were to ask a simple extractive question from a bunch of images we recommend to turn the reasoning off and setting the temperature to zero. But when you're going a little bit complicated like for example if I were to ask some kind of a comparative question uh you might want to turn the reasoning on and set the temperature to 6. This time I'm going to among the six images I provided from the PDF, I'm going to try asking which business unit had the most uh growth year-over-year. Let's see how it goes about uh pulling out the details. So I've set the temperature to 6 and increase the max tokens because we're going to generate a lot more thing tokens. We want to keep that budget a little bit high. And as you can see um it went ahead and pulled a bunch of details. It went to the data center section, went to the gaming section, professional V section and the automotive, all the four uh business units that Nvidia reports and uh looked at all the figures and came up with the conclusion that the answer is in fact automotive which had the highest year-over-year growth. That's great. So with improved reasoning capabilities um we were able to as I said we were able to increase multi-image understanding and that also gave us leeway to do amazing video understanding as well. Um so here I am going to provide a a video sample and try asking it to describe this particular video. Um this is a video from an omniverse rendering um showing you heart trees and multiple uh perspectives of the same 3D model. Let's go ahead and ask it to describe the video. So with respect to images um we um it does pretty well with any kind of document related data. uh be it tables uh charts or in general any kind of PDFs and with videos it's really um it is able to uh uh generate high high quality dense captions for any kind of uh unstructured videos that you could see in real world or it could be uh videos of uh people speaking um and people speaking or giving presentations just like this one. Let's look at the caption it generated. The video displays a 3D modeling software with small rustic wooden cabin. As you can see, it's it's it's capturing a lot of the details that you see in the video. Has a small porch area. Environment is covered in snow. And the video is likely a tutorial demonstration of the software's capabilities. I think uh I think that's a pretty good dense caption. You can use this for uh indexing videos within your um uh folders and be able to do be able to retrieve them based on the dense captions later on. >> All right, that's a quick demo of this model. Um Chris, do you want to take away with the introductions? >> Yeah, absolutely. Thank you, Andy, for showing us, you know, all of the the cool stuff we can do with NanoVL. Uh it's a really really awesome new model. Uh, so who are we? Well, uh, Zach, bring up the slides. Zach's the man behind the scenes making the show run smoothly. Uh, my name is Chris. Uh, I'm a product research engineer here at NVIDIA. And I'm joined, of course, by Annie, who is a developer advocate engineer at NVIDIA. Uh, and what we hope to do today is answer any questions you guys have about our newest set of released models uh, unveiled at GDC. Uh there are a number of uh of awesome models that we that we that we dropped uh over the last week and we'd love to be able to tell you guys all about them and uh and get you uh whatever whatever questions you have answered about these uh these models. Okay. So, uh the basic idea is uh I'm going to get my my screen up. It looks like uh Annie can can hear me again. What's up, Annie? Thanks for the awesome demo. And uh and I'll get my screen up here, Zach, and we'll we'll walk through uh kind of how you can look at these uh these new models. So through hugging face, uh we have our Nvidia or uh you can scroll down to collections and see some of our collections. And then of course you can continue to scroll down and see some of our models. We have a couple collections of note. Of course, we have our Nvidia Neotron V2 collection. This is going to include things like the brand new uh VL model, which as you can see can handle both pictures and uh video. Video of course being many pictures in a row. We also dropped our Neotron rag models. So these were actually already kind of out. Uh they they've been released before, but we've opened them up now. So you'll notice that you know beyond just uh having the the models available as a as a closed sourced option they're fully open. You can do uh you know you can play with the weights you can fine-tune them and they're they're incredible models right so for reranking and embedding uh these these two models specifically we also have of course OCR getting structured uh elements out of data anyway a huge suite of really powerful uh rag focused models were opened up and if you want to know how you can like use some of these things uh I'm going to point you to the Neotron uh uh GitHub the Neotron GitHub is where you can learn all things Neotron. Right now we have of course some examples and some usage cookbooks. This is like how to get it deployed in uh in VLM, right? Uh and use cases are like how do I use this to do something? And then of course we have uh coming soon we have some training recipes. So we're actually going to open source the code and put in this repo on exactly how we train these models so you can do it yourself. uh since we've not only opened the models but also uh the data and the recipes uh as well in the GitHub you can see of course any ideas you have drop them in the Neotron ideas portal uh where we monitor this to see uh what what people think is interesting uh and uh and what what kinds of things they expect their models to be able to do. Uh so that's that's what we've got going on. That was all last week. Dropped a bunch of models. Annie, what was your favorite model from the from the drop? And Zach, we can we can lower the screen down for a second while we while we chat about this. >> Can I be a little biased and say it was an Lotron Nano2VL? >> Yeah, of course you could be biased. Yeah, >> cuz come on, it it gives amazing video captions and it's something it's something that's mind-blowing that a small model like a 12B model could be able to generate such good dense captions. >> Yeah. And it's it's so good because it's based off the same backbone as the uh as the Nano V2 model which we've talked about a lot on these on these streams. Uh and so it's got the same kind of benefits uh you know that that we saw in in the in that model but extended into that new new domain. Uh my personal favorite models are the rag models. Uh we've we've used these a lot uh because they're great and they're finally open. Uh and openness is uh is huge. Obviously, it's a big uh it's a big, you know, part of the Neotron philosophy is having these open models. Uh Andy, what we'll do now is we'll take a couple questions. I see we've got a few questions uh in the in the chat and so we'll we'll talk about those. Uh number one, can I play with it in DGX Spark? Uh Andy, I don't know if you know the answer to this question. >> We haven't we haven't we're working on bringing the support to DGX Spark, but as of today, it does not unfortunately work on it. But we're working in the background to get the support on DGX Spark. >> So, uh soon is the is the answer and I imagine that we'll have a cookbook that will help you explore how to do that once we do and uh so you can just keep keep up with the with the the repository and we'll drop that when it's ready. Uh you can play with uh nano v2 on the spark. So there is a uh you can you can play with that that model uh run it you know have all kinds of fun with it. Uh there you go. Okay we'll go to the next question which is a great question uh because I have a lot of questions about this. So uh so is it a multimodal model then Annie? >> Of course it is but uh the modalities are the images video and text. So it is a multimodal model. So my question is we just talked about how this is built on a backbone of nano 2 uh you know LLM just the language model. So how did we take it from it's just the only modality is text. How did we take it from that to now we're handling video and images? >> So Nvidia has built uh an amazing vision encoder. It's state-of-the-art. um what what we do is we train this vision encoder to generate vision tokens that then the LLM layer tries to understand along with the textual tokens. So we're combining um a vision encoder along with a with an LLM backbone. So it's learning the textual tokens and the visual tokens together um um giving the model an understanding of both images and text at the same time. >> Yeah. So it it it it's like we kind of souped up the modalities by by adding uh adding an extra component that helps us translate some of this image data uh for the model. Is that is that the the correct understanding? >> Yes. >> Excellent. Excellent. And uh I see that we do video. What what goes into making video possible when we have uh you know obviously video is just a bunch of images in a row, right? So, what goes into making video possible versus images? Like, like how do we how do we know what what the start of the video is versus the end of the video, right? How do we how do we figure out that temporal relationship across the whole video? >> Well, honestly, um video itself is not going into the model as is. Um as as your understanding is right, video is a bunch of frames and a bunch of frames is being uh is is what is being processed. But the order in which each of the image goes in because at the end of the day the model is looking at uh the images as tokens. So the order in which the tokens flow is how the model was trained to understand a video input. Um so a bunch of frames going in and the order of tokens is how it's understanding the temporal nature of the uh nature of the images. And this is one interesting thing also um given the fact that you are sending in a video which means you're taking a bunch of frames and video I mean this model also comes with a uh an algorithm called EDS efficient video sampling that actually u where you don't have to I mean imagine if there is a 30 fps video and it's for 2 minutes you're going to have to process a bunch of frames right but EDS is u more intelligent it tries to understand where the uh density of the information is present within the video so if you have a lot of static frames you're going um reduce all of that to just like a couple of them which have more information within it. So you're reducing the number of u frames that you're processing. That means less number of tokens for the model to process and hence uh better speed. So this way we're also able to uh process videos much more efficiently and thanks to the EVS algorithm. >> Oh, that's so cool. So like if I have if I have like uh some security f cam footage which is quite static, right? I'm just looking at one spot. But then some crazy thing happens. This EVS is going to help make sure that I'm not really useless, totally unchanging, you know, video frames. Yes. It's going to help compress that down a little bit so that I'm I'm getting better performance. That's pretty cool. That's pretty cool. >> Uh, absolutely. We got a question about where do we run these? So, this is from uh Jagr Sharma. So, where do we run these? So, of course, uh, Nvidia GPUs, Annie, which ones should we run this on? >> Nvidia A100, Nvidia H100 are totally supported. In case you don't have access to those GPUs, we have build.envidia.com. We have an hosted API. So, feel free to go and play with it. >> So, it so it's optimized for those uh for those pieces of compute which you can you can rent from your favorite compute providers. Also, if you want to play with the model, I'll just get my screen back up. Uh, Zach, if you don't mind, uh, you can head over to our friends at Open Router uh, who uh, who have exposed an API for us. We've got some, uh, some compute providers that are they're helping this, uh, you know, helping this thing run. And you can see like we can ask the question u, you know, what do these charts show and pass in some charts and get a pretty, you know, solid answer about what, you know, what's going on here. The idea being that you know if you don't have the compute like Annie said you can go to build.invia.com or you can go to open router and either you know use the rate limited version through build.invia.com that's free or you can use uh uh you know a payer kind of model that you're used to through compute providers through open router. So uh you know it's not just uh it's not just the uh you know run it yourself. This is this is being hosted. Annie, one question I had is like, so this is this part's interesting to me, right? I have multiple pictures. I heard this model is exceptionally good at uh at reasoning across multiple images. Is that is that true? And like how should I think about when to because just like our other models, this model has reasoning on and off mode, right? H how do I think about when to run reasoning on versus when to run reasoning off with with with such a powerful vision language model? So the reasoning intuition still transfers across uh text models and multimodal models, right? So if you even though you turn on um you send in multiple images or you're sending only a single image or maybe you're not just sending any image, right? You want to use reasoning when the question is very complicated where you want the model to generate complex thoughts before it arrives the arrives at the answer. So if you are if your question even though has multiple images but requires you to do multi-step reasoning that's when we recommend turning on the reasoning but if you're like doing pure extractive um uh questions if you're doing a pure extractive based questioning on the images even though it's multiple images we recommend to turn the reasoning off. Does that make sense? >> Yeah, absolutely. Uh I mean you know where can we read about like some of these best practices like where do I find them in the wild? Right? If I'm uh >> well the >> go ahead. >> Well the notebook I have should capture uh most of it and we hosted that on the Neotron GitHub. So please go check it out. >> So we have our usage cookbooks. We have our Neotron Nano2. I'll zoom in a little bit here for you guys so you can see it. Uh Neotron Nano2VL and we've got our build general usage cookbook and our VLM cookbook. So these are going to help you understand uh how to use the model in the best ways as Annie has just described to us. Thank you very much for that Annie. We'll go to the next question now. And guys, please any questions you have, drop them in the chat. Uh what we're here to do today is answer your questions about all these dope new models that we that we drop. So anything you've got, put it in there. What's the best repo to learn how to do RL with Nemo? Uh well this is a great question and the idea is that we have in the invaded Nemo uh or we have the RL repository. So this is Nemo RL. Uh so this is this is the answer. This is the best place to go if you want to learn how to do uh RL with uh with Nemo. Uh there you go. All right. Excellent question, guys. That's right. There's just a repo for it. So uh we we got you covered there. We'll go we'll pop on down to the next question. Are the fine tuning parameters for this model different for images and text? And we'll we we're able to drop the screen now. Uh you know uh Andy I think this question is asking you know do we have to think about fine-tuning this model differently because there is images and text and if so how should we think about that? Well, well, the intuition of fine-tuning itself should not be very different. Uh, but that said, there is vision encoder involved and the parameter there are extra parameters that you'll have to tune uh than with respect to fine-tuning just the text model. We yet did not release a fine-tuning recipe, but that's coming soon as well. And uh uh just keep tuned in uh into the neatron repository. We will drop the fine-tuning recipe when that is done. >> Yeah. And this is kind of the idea, right? So, as we have any new guides or tutorials or any content that we have that we're creating to help people understand how to use these models, fine-tune these mod, whatever it is that you want to be able to do, this is the uh this is the the repo that you're going to want to go to is that Neotron repo. So, uh excellent stuff. Uh and great question from Prrenov there. Uh yeah, it's uh let's go to the next question. Yeah, hell yeah. Any tips for benchmarking multimodal accuracy across image, audio, and text with Neotron models? Quick caveat there. Uh while we do have audio specific models, we don't yet have an omni model that encompasses all of these uh modalities at the same time that that is part of the Neotron family. Uh but still an excellent question. There are we have plenty of of course Reva Parakeet models, right? We have tons of audio models, but just for this specific uh question, uh Neotron doesn't yet have a a audio solution, but it's an excellent question. How do we benchmark multimodal accuracy when we have so many modalities that that can play with each other? Any any tips that you've picked up over your experimentation for for this? >> Well, honestly, um there are multiple data sets that the community is following to understand u multimodal accuracy, right? And these data sets are particularly made to understand how best is it uh tying text within an image or how best is a textual question able to capture uh the information within an image. So um what we have rigorously test on was the OCR bench v2 although the word says OCR bench. Uh it doesn't really just do OCR but it has a bunch of other uh u kinds of um classifications underneath it like it would try to test the spatial understanding within an image. um it would try to um do visual question answering. So I I would say um OCR bench v2 or um the mmu um multimodal uh I can't recollect the entire thing with tripleMU uh which is also a good uh visual question answering benchmark that will help you understand the accuracies of these different multimodal models that are coming out and these are something that have popular uh where people are reporting their accuracy. So that'll be a great uh marker for you to understand. >> Yeah, I mean in in reality benchmarking uh models that have such impressive capabilities across a number of modalities is really going to come down your specific use case. Uh you know we can get these general benchmarks that Annie's talking about but it really depends on how you want to leverage these systems right. So like for computer use we have a set of benchmarks that you might uh you might think or care about right for like Annie is saying information extraction OCR we have uh some other sets of of benchmarks the the reality is is that because these are such impressive systems it's it's tough to have like a onesizefits-all uh you know set of benchmarks but once you have your kind of use case locked in then it it gets uh a lot easier to determine what you're going to actually need to benchmark Um though, and I and I say this uh with with with a great fondness, you're going to have to do a lot of work cooking up your own benchmarks for for very hypersp specific uh you know uh uh uh capabilities, which is fun and exciting. So there you go. We'll go to the next question. How is responsible AI rules considered for image and video model? This is an excellent question. Uh Annie, how how do we think about, you know, responsible AI now that we've kind of got image and video uh coming into the loop? >> Personal discretion. Personal discretion at best I guess. Um I think the models are uh of course trained to uh understand the safety uh within an image and a video and not and adhere from answering on such kind of uh modalities just like a text model. Um but yeah, at the end of the day, um as as a as a person who's building a solution, um or like a pipeline or like using these models, um it's just not the model that's responsible for the entirety of the uh safety within a within the uh thing, right? It's the entire system architecture where you embed these kind of responsible rules across your pipeline. Um but I would say uh we have industry standard safety u um you know baked into our training sets and it's pretty good at trying to understand that but beyond that I think it's a platform responsibility to understand the responsibility um or like the uh governance across >> yeah I mean and this is a great time to do a little bit of a plug right uh so we we actually have released uh another model in our batch of many models that we released at GDCDC which is to do with actually uh safety guarding. So a guardrail model and I this is something that we're really excited about which is this kind of uh you know this this idea that we can add guard rails to a to a system and that those guard rails help us to uh you know understand what it what is or isn't safe. uh and we have uh you know we've released the the text version of this which is in our uh you know in the in the re repositor or the hugging face uh organization uh great model uh and the the idea is again like Annie said it's a lot of discretion a lot of work goes into uh bias and safety training you know we have a a model card plus uh and that refers to this idea that you know we've thought about this safety uh we we thought about ethical use and you can read many details in the model card plus uh about you know where where where the shortcomings are and where the where the strengths are. Um but excellent question and again guard rails is something that's useful and and can help you ensure that your model stays on track. Before we go on to our next question from uh the the audience Annie I have a question. So, you're in you're at Nvidia and you you you work a lot on vision models. Like I I see you a lot you're on my team. I see you a lot being anme in this area. What what drew you to uh the vision side of the pipeline? Like what what draws you to uh giving these models like the ability to to to see both images and uh and videos? >> Okay. So personally I've been a computer vision girly back back when back before the LLMs and transformers took over the world. Um I I was natively working in computer vision and again let me also tell you personally I don't like to read books. I love to watch movies. Okay I believe in learning by through vision more than just reading at things. So I guess when the models um when transformers met met with vision and we are able to um you know natively understand images through text I think uh that's a that's an that's an amazing capability and I'm always drawn towards it and I'm excited that we're at this place and I there's a there's a longer road to go with respect to vision understanding. We're still we're still in the beginning just like how we you know LLM have matured. I think um the VLMs are just starting to get matured and uh there's a long way to go and we're going to see a lot more exciting things happening in the space just when we combine all the modalities. It's like how our brain takes in all the signals, right? >> Incredible answer. Vision vision from the from the start. Let's go. Thanks, Andy. Really, you know, it's it's fun to learn what what gets people down these paths. We'll take the next question from the audience, though. Let's uh let's keep the show rolling. Which Neotron features most reduce inference cost when deploying multimodal agents at scale? What a what a tea up. Uh yeah, Annie, I'll let you I'll let you take the the first uh first poke at this. I think uh the architecture um the especially the Mamba hybrid architecture uh that we introduced with nano tremendously brings the um um u or tremendously improves the throughput brings the latency down. Uh I think the architectural innovation is what brings it and then of course uh the optimizations we do uh on the hardware uh at NVIDIA uh that's something we do the best uh we know how to uh best optimize it for our hardware and that of course brings in the uh amazing speed that we uh that we have. Yeah. >> Exc could have answered it better myself. Uh 10 out of 10. uh you know one one thing that's interesting about this model as well as nano is they're not just available in in you know BF-16 or FP8 uh they're they're also available in NVFP4 which is uh quite a uh you know quite a uh an extreme quantization right all the way down to FP4 uh but optimized to run in that hardware that uh that that uh you know architecture the idea being that you know you can really leverage these smaller faster models to help bring those inference costs down. Uh and we we like to, you know, we like to say uh faster smart faster models are smarter models, right? So, uh the whole stack is set up to to give this uh to to do this. Uh one of the things you might have heard Jensen talk about in the in the keynote at GDCDC if you if you're able to turn it tune into that is this idea of extreme code design, right? So the models are designed with the chips are designed with the data is designed with the models is designed with the chips, right? And the idea being that building this this nice uh you know uh you know kind ecosystem helps us make sure that we're we're getting the we're squeezing the most juice out of the orange that you possibly can at all steps of the pipeline. Uh and of course got to shout out our friends at uh Nvidia Dynamo. Uh Dynamo is an inference platform that uh that helps you not have to think about a lot of this stuff. Yeah. Uh we'll go to the next question though. How do you handle scalability when multiple specialized agents need to collaborate in real time? This is an extension of the other question, right? So basically like how do we how do we think about when we have many specialized agents because our we're we're we're on the record in our blogs in our in our vlogs in our uh podcast Ray is talking about a specialized many specialized agents is something that's uh that that that's something we're excited about. So how do we how do we handle that that interaction? I think uh you know what Annie said earlier thinking about the uh thinking about the architecture making the architecture as as fast as possible as efficient as possible everything we're doing is is helping to do this uh this uh this idea of making these systems scalable any any anything else to to add here >> I think I think um one more thing with uh GPUs in general at scale is where we excel so more more the number of requests, better the throughput. So when we're talking about specialized agents, you're really getting the juice out of your GPU. So at scalability, at scale, we're um I think that's where we shine. >> Doesn't get better than that. Of course, we'll go to the next question. Where's the cookbook located again? Zach, can you throw up that uh that that that link again? It's at the Neotron repository. If you can bring up my screen for a second as well. Neotron repository. Okay, so this is Invaded Nemo/neotron. You can see that here. And uh you're going to go to the usage cookbook, and you're going to look for the model you care about, and you're going to click it here. If you're if you're wondering which model you might care about, you can scroll down. You can actually see a little table of models, what they do, what they're meant for, and then you can click on the one you want, and that's going to take you to the uh the the the model. Uh and then if you obviously in this case, it didn't. [laughter] Then uh if you want to go to the cookbook uh you'll be able to to follow a link as well. That's the idea. And uh I know who's about to submit a PR after this to update the reference to the model and his name rhymes with miss. Okay. Uh we'll go to the next question. Do you have any courses that can teach us anything? Ah you know what? I think we may just have some uh some learning paths for you guys that I know uh I think uh uh our behind the scenes legend uh Rebecca has has uh put a link to and it's on the screen right now. Uh basically the answer to your question is yes we do. Uh we have a number of different uh paths that can help you get started. It is overwhelming. I saw some comments. It's overwhelming. There's a lot where do I start? There's so much information and we're going to recommend you guys start at these learning paths uh specific to uh to to Agentic AI for for Neotron. It's going to help you get started, help you teach the core concepts you need uh to to get going. That's the uh that's the one uh excellent question though. Excellent question. Also ask questions in office hours. We hold these all the time. Uh so if you're ever worried about where to get started, where to go next, please just come in and ask. We'll we'll we'll help you as best we can. Uh, we'll go to the next question. Let's go. How do specialized agents built with Nemo guardrails handle cross agency or cross agent policy enforcement? Can you cascade safety rules across agent boundaries? What a question. Uh, Annie, I'll let you take the first crack with this one. Uh, any thoughts here? >> Cross agent policy enforcement. Um from what I understand, Nemo guardrails allows you to uh tune your guardrails for each of the policies separately. So um as long as each of the models each of the models have certain policies that you are um you're already setting for like each model has to handle XYZ set of policies, I think you're you're all handling that at platform level. I think this is more um more on how you design the agent architecture at the end of the day. um Nemo guardrails is going to do it at an individual uh policy level but not really help you build that entirety the entirety of the um architecture but yeah at a singular level at the end of the day it's up to you on how you want to design that cross policy enforcement >> if you want to add >> I mean hitting the nail on the head here uh the idea is that you're we want to think of agents as much as possible as small isolated components and we want to make sure that we're enforcing at component level. Uh if we're doing any kind of cross boundary uh you know uh enforcement, then we're we're going to want to make sure that we have some specific guardrail or policy set up for that. Uh but it it comes down to your architecting. Uh less so the specific technology, more so how you're tying everything together. Uh you can use something like NAT neo agent toolkit to help you do this. uh but uh the the the long and short of this is an architecture specific uh you know uh implementation. So great question though. Excellent question. Really thinking about making safe scalable agents. So you'll love to see that. We'll go to the next question. What's the best pipeline approach to deploy this to prod? Oh, you know what? I'm just going to say it. It's nim. Uh we have a nim. That's all it does. Uh Annie, any anything you want to add? >> I think I think so, too. Head over to build.envidia.com and I guess you'll see a button to deploy it to Nim. >> Basically, Nim is our container that does this for you. So, we're we're that's what we're going to suggest. Uh if you want to customize it or play with it, you can do that with our other frameworks. Uh but Nim is the the way to go. And you can start with build.vidia.com video.com as Annie said and then transition and then when you're good and ready. Uh, excellent stuff. Thank you Annie. We'll go to the next question. What tool or APIs you recommend for tracing and debugging agent decision steps across vision language and ash components? You guys are you're you're teeing us up too well today. Andy, where where do they go? Where do they get started? >> Honestly, tools APIs are all very personal to what kind of um what kind of requirements uh you have. Of course, we do not want to recommend one over the other, but we would say get started with uh the Nemo agent toolkit that um Nvidia uh offers. It gives you it gives you good tooling for in general uh building your agent, getting started with it, and then of course the traceability and all of these are features that come with it. So, um again, this if if you get started with Nemo agent toolkit, you should be able to also connect with multiple other toolkits that are in the ecosystem. So it's a good starting place uh because it has gotten good templates and all the features that you'd require to get started with building agents and you can obviously move from there if you were to use any other framework. >> Yeah, I mean exactly. So Nemo Nemo engine toolkit is a place to get started with this whether you're using arise whether you're using uh you know whatever you're using doesn't matter lang lang chain crew uh llama index every everything right uh Nemo agent toolkit is going to help you glue these pieces together in order to to build the best combination uh for you and slotting into your existing uh tech stack. Thank you so much for that question. We love talking about Nemo Agent Toolkit. Okay, next question. When do you This This is where we're going to wrap it for today, guys, because we're we're a little bit over time, but we love chatting with you guys. Uh uh where when do you guys host the office hours? So, we're actually we're here every single Tuesday mostly. I mean, holiday is, you know, accepted, but uh we're here on Tuesdays at 11:00 a.m. uh Pacific Standard Time. Now, hope everybody enjoyed their extra hour of the weekend. Uh, and we're we're always happy to answer questions. However, that's not the only place that you can find us. Segue to some slides uh where we uh you can reach us outside. So, hey, there we are again. Uh if you want to learn more about some of what we talked about today, check out the Aentic AI uh day at GDC DC, which is now available uh on demand, meaning you can watch the the recordings uh to deeper dive into some of these topics. Uh we'll go to the next slide. Uh if you have questions that are so pressing you cannot wait till Tuesday or you just want to chat with us, head on over to discord.gg/invidi developer.nematron Neatron models is the channel that you'll find me in specifically all the time. Just ping me and as well as a number of the other members of our team. Please ask us questions whenever you want. There's no uh you know there's no limit on when you can ask questions. Of course, you can always come on a Tuesday and pop your question in the chat, but uh you know this is a place you can ask async. And we'll go to the next slide. Uh you can also just email us directly at communityinvidia.com. That's right. Email's open. Please let us know what you got going on and what questions you have. Share any awesome projects you've done with us. Uh we we really want to make sure that we're we're we're building a community here. We're we're you know, I'm a developer. Annie's a developer. This is what we do. This is what we love. Uh and we we love to do it with you guys. Uh and uh there you go. Yeah, that's us. Well, guys, thank you so much for joining Neotron Labs live stream. We're going to be here uh next week. We're going to have some friends on from uh I heard a pretty cool agent framework called Crew AI. Uh so come on down and join us for that one. That's going to be a blasty blast. Uh thanks so much for turning in and uh we'll we'll see you in the next one. Thank you. Bye. >> Okay.
Original Description
Join NVIDIA Nemotron experts live as we dive into what’s new for developers following GTC DC. We’ll explore how the latest advancements in agentic AI are making it easier to build intelligent, multimodal systems that see, reason, and act with greater precision and safety.
In this livestream, we’ll:
- Unpack the Nemotron open models, open datasets and techniques that help you create domain-specialized AI agents ready for production.
- Discover how advanced reasoning, multimodal perception, and retrieval systems can drive the next wave of AI innovation.
We’d love to hear your thoughts and experiences as you experiment with the Nemotron family of models—your feedback helps us improve performance, usability, and developer workflows across the ecosystem.
Tune in for a live walk-through and discussion designed to help developers push the boundaries of what AI agents can achieve—efficiently, safely, and at scale.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from NVIDIA Developer · NVIDIA Developer · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Ray Tracing Essentials Part 2: Rasterization versus Ray Tracing
NVIDIA Developer
Ray Tracing Essentials Part 3: Ray Tracing Hardware
NVIDIA Developer
Ray Tracing Essentials Part 4: The Ray Tracing Pipeline
NVIDIA Developer
NsightGraphics 2020 2 Release Spotlight
NVIDIA Developer
Ray Tracing Essentials Part 5: Ray Tracing Effects
NVIDIA Developer
Ray Tracing Essentials Part 6: The Rendering Equation
NVIDIA Developer
Ray Tracing Essentials Part 7: Denoising for Ray Tracing
NVIDIA Developer
Spatiotemporal Importance Resampling for Many-Light Ray Tracing (ReSTIR)
NVIDIA Developer
Announcing Cloud-Native Support for Jetson Platform
NVIDIA Developer
JetsonTV: Build your next project with NVIDIA Jetson
NVIDIA Developer
Nsight Compute Feature Spotlight: Roofline Analysis, Asynchronous Copy, Sparse Data Compression
NVIDIA Developer
Nsight Systems Feature Spotlight: OpenMP
NVIDIA Developer
Isaac Sim 2020: Deep Dive
NVIDIA Developer
NVIDIA Jetson: Enabling AI-Powered Autonomous Machines at Scale
NVIDIA Developer
NVIDIA Tools to Train, Build, and Deploy Intelligent Vision Applications at the Edge
NVIDIA Developer
Jetson Xavier NX Developer Kit: The Next Leap in Edge Computing
NVIDIA Developer
Synthesizing High-Resolution Images with StyleGAN2
NVIDIA Developer
NVIDIA Robotics: Isaac SDK and Sim 2020.1
NVIDIA Developer
Accelerating COVID-19 Research with GPUs
NVIDIA Developer
Visualizing 150 Terabytes of Data
NVIDIA Developer
Boosting Performance and Utilization with Multi-Instance GPU
NVIDIA Developer
Running Multiple Workloads on a Single A100 GPU
NVIDIA Developer
NVIDIA Nsight Feature Spotlight: GPU Trace
NVIDIA Developer
Spark 3 Demo: Comparing Performance of GPUs vs. CPUs
NVIDIA Developer
NVIDIA Jetson Nano Wins Edge AI and Vision Alliance Award
NVIDIA Developer
NVIDIA IndeX on Google Cloud Platform Marketplace
NVIDIA Developer
DeepStream SDK: Best practices for performance optimization
NVIDIA Developer
Efficiently Deploying GPU Accelerated 5G CloudRAN for Edge AI Inferencing
NVIDIA Developer
NVIDIA PhysicsNeMo - Accelerating Scientific & Engineering Simulation Workflows with AI
NVIDIA Developer
NVIDIA Deep Learning Institute Instructor-Led Training Available Remotely
NVIDIA Developer
Advancing AR Glasses
NVIDIA Developer
Blender Cycles: RTX On
NVIDIA Developer
Real-Time GPU-Accelerated Data Analytics of 250 million Flight Data Records of 737 Max grounding
NVIDIA Developer
Assessing Property Damage with AI
NVIDIA Developer
RAPIDS: GPU-Accelerated Data Analytics & Machine Learning
NVIDIA Developer
DaVinci Resolve Turns RTX On
NVIDIA Developer
RAPIDS with Plotly Dash : GPU-Accelerated Census 2010 Visualization
NVIDIA Developer
NVIDIA IndeX for arivis5D Cloud Platform
NVIDIA Developer
NVIDIA Backchannel: Behind the Scenes of Marbles at Night RTX
NVIDIA Developer
NVIDIA Backchannel: Sneak Peek into Marbles RTX in Omniverse
NVIDIA Developer
How to Create "Paint" in Substance Painter
NVIDIA Developer
Accelerate AI development for Computer Vision on the NVIDIA Jetson with alwaysAI
NVIDIA Developer
Securing Next Generation Apps over VMware Cloud Foundation with Bluefield-2 DPU
NVIDIA Developer
Accelerated Data Centers with NVIDIA and VMware
NVIDIA Developer
GPU-Accelerated Motion Blur in Blender Cycles
NVIDIA Developer
NVIDIA Clara Guardian Virtual Patient Assistant
NVIDIA Developer
Revolutionizing Supercomputing with NVIDIA UFM Cyber-AI
NVIDIA Developer
Inventing Virtual Meetings of Tomorrow with NVIDIA AI Research
NVIDIA Developer
Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
NVIDIA Developer
Getting started with Jetson Nano 2GB Developer Kit
NVIDIA Developer
NVIDIA Jetson Developer Community AI Projects
NVIDIA Developer
Open-source projects on NVIDIA Jetson Nano 2GB Developer Kit
NVIDIA Developer
Real-Time Ray Tracing with Project Lavina
NVIDIA Developer
Jetson AI Fundamentals - S1E2 - Hello Camera
NVIDIA Developer
Develop Optimized Conversational AI Models with NVIDIA NeMo on DGX A100
NVIDIA Developer
Jetson AI Fundamentals - S1E4 - Image Regression Project
NVIDIA Developer
Jetson AI Fundamentals - S2E1 - JetBot Intro and Hardware
NVIDIA Developer
Jetson AI Fundamentals - S2E2 - JetBot Software Setup
NVIDIA Developer
Jetson AI Fundamentals - S1E1 - First Time Setup with JetPack
NVIDIA Developer
Jetson AI Fundamentals - S1E3 - Image Classification Project
NVIDIA Developer
More on: Agent Foundations
View skill →Related Reads
📰
📰
📰
📰
Can we truly understand maternal mental health, or are we only measuring it?
Medium · Deep Learning
How Neuro Linguistic Programming Helps with Anxiety and Stress Management
Medium · NLP
Your Attitude Is the Lens — Change It and the Whole World Looks Different
Medium · Deep Learning
Mental Models Don’t Just Shape What We Know. They Shape What We Can Learn.
Medium · Deep Learning
🎓
Tutor Explanation
DeepCamp AI