Build Custom Large-Scale Generative AI Models | NVIDIA GTC

NVIDIA Developer · Intermediate ·🧠 Large Language Models ·3mo ago

Key Takeaways

Building custom large-scale generative AI models using NVIDIA NeMo and Adobe Firefly Foundry

Full Transcript

All right, hi everybody. >> [applause] >> Um, very glad to be here. Thank you all. Thank you Ingrid and Vidya for inviting me here to speak. We've been partners with Nvidia for I don't know many, many years going back more than a decade. That started with us working very closely with them on optimizing all of our flagship desktop tools to run on their client side hardware. And as we've turned into this GenAI area, we now work very closely with them to really figure out how to optimize the insane task that we decided to take on of building our own frontier models from scratch and then serving them at scale with all of our customers to be able to provide lots of wonderful GenAI value. So thank you guys for for doing that. Um Okay, so let's have a talk about how the talk's going to go. If you're here, I'm assuming it's because you're interested in possibly building models yourself, building large-scale models because that's really what we're going to focus on here. The relative scale of kind of where we're going to go here technically is probably middle of the road. I'm as Ingrid said, I'm CTO for Adobe which means I know just enough about the technical depths here to really be dangerous. So I'm going to shoot for that middle of the road. If you've never done any training, you should be able to do this. If you've done some training there's probably things in here for you to learn about what it means to do that at scale. If we do end up having time for questions and there's any you guys want to get into the weeds, I have Marion here from our team who knows way more than I do about all of this will be able to help us as well. >> [snorts] >> Okay, so the way I'm going to frame this is I'm going to talk about the journey we went on starting from why we decided to take on this crazy task of building large-scale models, how we made that decision, what the pros and cons were and then of course what we learned along the way and all the battle scars we ran into. So hopefully you guys will pick up some lessons and be able to be able to learn from them and avoid some of them when you jump in. So I'm going to jump back here first about 3 years ago. Right, so this was I don't know it was like late 2022. ChatGPT wasn't on the market yet. Nobody had you know that that mind-blowing change to the GPT world hadn't happened. But what we were seeing was a lot of media models, you know, really starting to hit the stride of they could actually provide real value in the research labs. And you know, mid to late 2022 is when we saw the first models actually coming out commercially that could provide real results for end users to use. released DALL-E 2, Stable Diffusion came along and blew everybody's mind with an open source model. And at Adobe, right, our job is taking the latest creative technology and and putting it in the hands of creatives and to accelerate their craft and let them you know, create more amazing things. And so you know, we very quickly looked at this and said, how do we get this into our products? We went and looked at the models out there and we thought about what kind of features we we could build and we very quickly realized there were some problems. The problems were while these models were amazing and could do amazing things, they weren't going to meet the needs, the specific requirements of our customers and so we might have to figure out how to go above and beyond that. I'll give you two specific examples. First is control. Right, if you think back to that time we had you know, what we used to affectionately call prompt roulette. You type a prompt into the you know, into the the thing and say like I want dogs playing poker or some you know, internet cat picture and it would generate this amazing picture for you. And if you didn't like what you got, you know, you change the prompt a little bit and you'd get a new picture really fast but it's probably completely different from the first picture you got. For our customers, that's a non-starter. Right, these are highly trained creatives. They bring their taste and their vision and that's you know, as much as the technical ability, that's what they bring to the table. And so putting this technology in their hand where they couldn't iterate and control down to the fine details was not going to be useful to them. So we knew we needed models that could really push the bounds on control. The other one was really important turned out and we found as soon as we started talking to customers, especially our large enterprises which was all of the legal and ethical and regulatory questions about these models. Now again, if you think back there, you know, there was a lot of public discussion going about about about the some would say dumpster fire of concerns around where these models were trained, where the data was taken from, what rights do people have to train on it. Are you going to get in trouble for using a model that took you know, data from some website? Are you going to get in trouble for using a model that can completely mimic the style of a well-known artist? Do you own the content that you're putting out there? Are you going to suddenly get shut down in Europe because you're using this stuff? Every one of our customers basically was telling us they couldn't touch this stuff with a 10-foot pole. Not that they didn't want to but their legal department was saying no way. And so we had to figure out how can we actually make this stuff available in a way that they feel comfortable adopting to get moving with. And that really came down to us going all the way back to the beginning starting with a licensed clean data set that we verified with human moderation had no copyright in it, no IP that was going to get somebody in trouble. We could trace the providence of every asset that we used to train on and we could give that to our customers with very high guarantees that don't worry, this is safe to use. Right, so that stuff was critical. Right, and we knew this was what what the market needed and we also knew it was a place we could differentiate. But that put us in a difficult position and it's one that you know, if you're in the tech industry, you've you've faced many times a year probably which is basically build it or buy it. Right, we could [snorts] grab these off-the-shelf models which were these you know, marvels of technical wizardry and large amounts of manpower and dollars going into building them and we could hit the ground running but it wouldn't really meet our customers' needs. Right, we'd be a lot of compromises and it would compromise who could actually use it. Or we could roll up our sleeves and you know, dive in and say we're going to build it ourselves. We had access to a lot of licensed data so we knew we could do that. We had brilliant people. Adobe has had we have a research team of you know, over 200 PhDs who have been doing you know, world-class leading computer vision research and AI research for over a decade. So we thought we knew how to do it. Um, it was just a question if we wanted to take it on. And you know, this is obviously the first question for you to ask as a technologist. I always like to build things but for something of this scale with this level of investment, really have to take a step back and ask what is the benefit that we think we're going to get from this? Is this is what we can build with this differentiated from what we can get off the shelf? Um, and does that differentiation justify the investment that we're going to have to put into it? And that investment is large as you'll see in a minute. So we did that calculation and we said yep, yep, definitely important. We think we should do it. We think you know, how hard could it be? Ha ha ha, right? Actually you know, this was a little naive of us. The answer is actually not that hard. You know, like I said, we have very smart people at Adobe. We've been doing AI in a while and so we knew how to train models. This was just kind of like throwing a couple more GPUs at it, right? So let's talk about how that works. Um the the basic idea here is again, if you've trained a model on your laptop, it's it's really no different from what you've done. This is a you know, a picture of a of a typical training loop. Right, simple but in some ways it is this simple at least from the 10,000 foot view. What you do is you go gather gather a bunch of data and you set up this loop and you get a GPU and you say okay, I'm going to load the data off disk and then I'm going to do some prep work. I'm going to figure out how do I you know, transform the data so that it's consumable by the GPU. And then once I get it there, I need to turn it into something that is consumable by the model. Right, I have to turn it into an embedding. Um, and then I hand it off to the model and say okay, great. Process this. Look at what you did. Figure out how far you off. Make an adjustment to the model so it's a little bit closer to the right answer and then just repeat that thousands or millions or hundreds of millions of time and you've got a nicely trained model. And then every once in a while just to be safe just like you do with anything you're doing, you know, you check it back in. GitHub, you know, save a copy in your on your Mac, whatever it is. We take checkpoints every once in a while along the way and we just grab the model as it is and we say okay, we're going to put that back in storage so if something goes wrong and things go wrong all the time for all sorts of reasons, you know, this is as much an art as a science. We can always go back to one of those checkpoints and pick up again and you know, try going a different direction. Right, so that was the basic model. Again, this is what you do in your laptop. This is what we were going to do in the cloud. We just needed a few more GPUs and so we said great, let's go ahead and get to work and do it. So we built a very basic training stack. Right, a little bit more complicated than what you did doing your laptop because you know, we had to do a couple do a couple things to do this. So first of all, we went and got a bunch of couple thousand high-end machines that had the best and you know, brightest, latest Nvidia GPUs in them. I think I think they were A100s at the time. I don't remember. Um and then we got some storage. You know, we had our data and the data in this case, you know, we had petabytes of data and it's growing all the time. Mix of images and videos and other assets, low res, high res, long videos, short videos, different codecs and formats. So we had all this data. We needed to put it somewhere. So we used a relatively standard distributed you know, storage infrastructure on on our our provider at the time was AWS. So we you know, just used some of their infrastructure S3 to put our data up there. And then we grabbed some Python code to glue it together and and run that run that loop. Let me talk about that Python code for a second. If you've done training, you know what PyTorch is. For those of you haven't, this is the kind of one of the industry standards of how you write you know, complex model manipulation and training code to you know, to to run on the GPU so you don't have to get deep into the weeds. It's great. It's a very powerful library. Even that we said we don't think we need all that power. There's this library that exists called PyTorch Lightning which is sort of a level up above PyTorch that says hey, if you're going to do a standard training loop, that's easy. Just you know, use this. We'll get take care of all take care of all the boilerplate code for you and and we'll you know, you can just kind of hit a button and get it going. So we said great, that sounds like what we're doing. Set up that loop with PyTorch Lightning and and we were about ready to get go. We did have to make one decision at this point, which was, okay, we got petabytes of data, we got thousands of GPUs, we have to figure out somehow how do we slice up that data, how do we slice up the GPUs to take advantage of the fact that it's not just one, but all these guys are going to hopefully accelerate training by working together. There's a lot of decisions you can make here. Um, we took the most straightforward one. We use something called data parallelism. What it basically is is we take our, you know, let's say we had a thousand GPUs, we take our petabytes of data and we divide it up into a thousand chunks, um, and we ship each one of those off to GPU with a copy of the model. We have it do its training run and once it's been through all the data and it's, you know, updated the model, we, you know, have another machine which collects all those changes back together, you know, collates the results, takes all the work they did and figures out how to merge it, and great, now we have a new checkpoint. Fantastic. Okay, so we did this, we set it up, we got it running, and in a very short time we actually had a model. And it was a great model, and you know, this was our first breakthrough of like, oh my god, you know, this is not rocket science, right? It is rocket science, but we had rocket scientists, and and we could do this, right? We had a model that could create images. Um, you know, from a simple prompt. This was 2022, so yeah, sure, the people had seven fingers and, you know, all that kind of stuff, but that was the state of the art at the time and we were pretty excited about it. Um, but seven fingers isn't really going to work for us, so we knew we had to get better and we had to get larger and we had to invest to make this thing better. We were ready to do that, but we noticed a problem. And the problem was that we were not getting anywhere near as much return on the investment of those GPUs as we thought we should have been. You know, we worked very closely with Nvidia, we know how powerful these things were, that's not what we were seeing. And if we were going to pour millions more dollars into building up our fleet and building this thing out and going much, you know, bigger with our training runs, we had to solve this problem. So we whipped out our profiler and we started looking at the GPUs. And we saw something disturbing. And let me show you an example here. So what you're looking at is a simplified version of what our profiler shows us. This is, um, two different GPUs that we're looking at. The top one is one of is sort of one of the standard training GPUs. So we got a thousand of these running, whatever it is, and they're all sitting there taking their data and processing and updating the model and running as hard as I can. All those nice colored lines, that is, you know, evidence of the GPU thinking hard and doing all this amazing work that, you know, that none of us can imagine even how it works, except the rocket scientists at at Nvidia. Um, the bottom line is a different GPU. That's kind of a a special manager machine. And I, you know, I said earlier, the way this training at scale works is all of the all the machines do their work and then there's a sort of a manager that does this gather phase, where it collects all the updated models back and it collates them together. You know, it's like the end of a test where, you know, like you finish your exam and and your teacher says, okay, go, you know, take five, go go have have a coffee, cigarette, something, I'm going to I'm going to grade and and get back to you guys, right? Just to pull them together, balance everybody out against each other, grade the curve, you know, whatever, curve the grades, etc. And you'll notice something, which those blocks of work that that bottom GPU is doing line up very well with that big blank section in the top one. That blank section is really big. It's about, you know, if this is time, that's about two-thirds of the time the GPU could be working. It's sitting there smoking a cigarette, doing nothing. Because that manager has to do a lot of work, right? The the job of that manager, taking all of those models and bringing them back together and doing that merging, um, that turns out to be really hard and expensive. Uh, and the more GPUs you have, the harder and more expensive it gets. Um, so that was a big problem, right? If we were putting a million dollars into training, you know, that was $600,000 that we were burning away on those GPUs sitting and doing nothing, right? So we were getting a third of the return on that, um, that we needed. Um, that is a big problem, it's a big problem, you know, for our ability to get the results we want, it's certainly a big problem when I go in front of our board and explain explain to them why they have to spend a lot of money on doing this, like telling them that we're getting 30 cents on the dollar is not a fun conversation. So, um, we had to fix this. And so we, you know, again, smart people, we rolled up our sleeves, we dug in to figure out the problems were. The problem is this is only one example, this problem with with data parallelism. And by the way, this example is is is real. I think, you know, my rocket scientists tell me these days that like if you're going to do a training run with more than like 16 GPUs running in parallel, data parallelism is going to run into these kind of problems. So, nice naive approach, we had to do something better. And there were many, many other places when we look around that entire training loop where we were seeing bottlenecks left and right. So, you know, we've got this distributed storage. Um, loading petabytes out of data over the network, uh, sorry, petabytes of data out of memory, that takes a long time and your GPUs are sitting doing nothing. Saving that big giant checkpoint, that huge model, you know, billions of parameters back into storage, again, huge times loading it, saving it back in. Again, you know, this is really insurance. We like most of the check checkpoints we'll never need, and so all this time is being spent just in case and those GPUs are sitting idle. In that core loop, right? Taking all that data and then loading it up, there's actually a lot of bottlenecks in there. So, um, just preparing the the data for for use on the GPU, right? That's all CPU work and there's memory bottlenecks there that prevent the machines from running fast enough. Um, getting into the into there and and loading up and using your VRAM effectively, using your computer effectively, lots of bottlenecks in there as well. Um, really interesting one there. Um, these GPUs, not only are they waiting for that manager at the end, like I showed, they're actually waiting for each other as well, which shouldn't happen because, you know, they're they're running fairly independently, that's the whole point of of working at scale. Um, but, you know, let's go back to that, uh, not very good, you know, teacher test example. A lot of you guys are probably the smartest people in your classes, you know, your test could come along and you, you know, do it and you finish in 10 minutes. And what do you do then? You sit around for the next 50 minutes waiting for the slowest kid in the class to finish before the teacher says, okay, everybody can get up and go and collect the test. That's exactly what happens with these GPUs. So remember, our data is very heterogeneous, right? We had videos, images, small res, high res, we had some that had very expensive codecs that took a lot of work to to, uh, process, some which were just raw data that didn't take any processing. And so we were naively slicing up our data in a way that meant we could not predict in any way, shape or form, or we weren't predicting, how much time each GPU was going to take. Some of them would go really fast and some of them, because they got the hard work, right? They were the unlucky guys who got the big expensive movies, they would take way longer than the average because that's the way our data kind of fell out. So, across all of these places, we realized, hey, just doing that kind of out of the box thing where we grab some nice libraries off the off the shelf and we we kind of throw them together and create this loop, it wasn't going to work, right? We were not going to be able to, you know, scale our usage enough to really be able to justify all the money that we're spending on doing this, even though we knew there was a pot at the end of the rainbow there that, you know, we really were going to create something different that was going to be meaningful for our customers. Um, hey, I had lost my, uh, my speaker notes, Mom. I don't know if something's going on there. Um, I will I will charge ahead, but, uh, who knows what I'm going to say at this point. Okay. Um, while they're working on it. So, uh, so again, you know, this was an investment for us and it was a big investment of time and learning and, you know, technology. We had to build a bunch of things ourselves. We had people who really had to, you know, get a crash course in this stuff. Um, so this is, you know, another check in sort of where our our training loop evolved to. Same basic structure you saw here, but there's a bunch of, you know, uh, accelerations and, you know, instead of deeper decisions and optimizations we put in. I'll touch on a few of them, you know, look at the top left and right. Loading data out of of memory and storing it back, sorry, out of storage and storing it back into storage. Loading [snorts] it out, okay, we just we that was theoretically not that hard. We we just created a specialized balanced data loader, which instead of blindly slicing up our data, it actually, you know, we built knowledge into it of what the data was, what the cost of the different data was, what kind of processing was needed, and we actually built an algorithm that sort of predicted, okay, how much, you know, how how do these inputs affect the workload that a GPU is going to do, and tried to balance it to figure out, you know, how we can leverage that out so their work is distributed pretty evenly between them. Um, on the storing side, you know, like I said, we were storing this big completed model, you know, every five training runs or something, I forget the exact number, and most of that we will never use again, it's insurance. And so it's kind of a shame that we're doing all that. And, you know, our team had the bright idea of rather than, you know, storing a beautiful completed model that an end user could use, even though we were never going to do that, we could store actually some of the pieces earlier in the process, some of the the data coming off those individual GPUs. If we stored that back into the distributed, uh, storage, we can, you know, you take advantage of the parallelism of the storage, um, and, uh, and, you know, sort of save it a lot faster. Yeah, if we had to go back to a checkpoint, it would slow down because we'd have to rebuild the model from those pieces, but, you know, that happens rarely, you know, that we can defer that work until it's needed if it's ever needed. Um, in the middle here, you see it says FSDP plus TP. This is our new, uh, parallelism approach. Stands for fully sharded data parallelism and tensor parallelism. Um, interesting topics there, go read about them if you want. The the long and the short of them is, rather than just slicing up the data and giving a complete copy of the model to every, um, GPU, we can actually slice up essentially the model itself, you know, various directions, slice it up by layer, slice slice it up by quadrant, you know, again, sort of read up for the details, but >> [snorts] >> and then we hand data and just a piece of the model off to every one of our thousand GPUs. And what that means is the GPUs I think end up doing a little bit of work, I'm actually not sure, but more the act of collecting those and gluing them back together is way, way simpler. It's like, you know, working on a puzzle with somebody. If you're all trying to like work on every other piece of the puzzle and then you got to merge them, like that that's going to take forever. But if you carve it up into blocks and you say, "You work on that piece and I'll work on this piece." Just easy job of gluing them back together. So, huge improvement on the performance of that gather step by being more advanced about how we slice up our data and then how we slice up the model for distributing among these. Um couple of the things actually I didn't talk about this in the previous one, but if you look in the lower right there, it says EFA. Um our first version of the training loop ran on, you know, standard machines, standard um standard ethernet on the backbone, which was wonderful except you're shipping petabytes of data around between thousands of GPUs millions and millions of times and that network speed becomes really, really important. Um turns out ethernet is not really up to the task of what we were doing. And so, there are fantastic technologies out there that are evolutions of the adapter and and the networking stack um that are very, very high performance and designed for these kinds of use cases. Uh Nvidia has one called NVLink, which is a fantastic piece of technology. Because we were running in AWS, you know, they had standardized on a different one called EFA. So, you know, we plugged that in um and uh and you know, that that was a huge bottleneck we removed. In the middle here, you know, you see it's not PyTorch Lightning anymore. Now it says PyTorch. Again, the off-the-shelf libraries are great if that's what you need, but if you need off-the-shelf if you're going to do it with off-the-shelf libraries, there's a good chance there's actually a model out there that you should be using rather than just using libraries and building it from scratch. So, probably you're going to have to go a click down if you're going to actually build this yourself. And so, we had to go down and you know, learn the intricacies of PyTorch, which we knew, again, we were using it for other things. Uh and and really optimize the actual training process itself. Um and we did that. We actually This is where we paired up with Nvidia and we spent um a lot of time really looking at how each operation that we were running runs, you know, exactly on the GPUs we were using and and just tune the hell out of those things within an inch of their life. Um you know, that that clearly we break that up into two buckets. A bunch of the stuff we did with them is is stuff that could be generalized uh to, you know, other use cases. And so, our friends in Nvidia, you know, pulled those back and put them back in their um cuDNN library, I think I'm pronouncing it right, which is, you know, the library that essentially implements these operations on their GPUs. Um and so, those are all, you know, now part of the libraries, I think, for everybody to benefit from. And there were definitely some that were very specific to our use case and our training loop. And so, we were also writing custom kernels there um to be able to just, you know, find the right optimizations for what we were doing. One last um uh optimization I'll talk about um thousands of GPUs, I don't know how many hard drives running up there. Uh these are all high-end consumer grade machines, right? They've got these beefy chips in front of them, but you know, a lot of the components are the kinds of things that, you know, we're running on our desktops. Um when you get thousands of those running, you know, at full speed, full heat, you know, 24 hours a day, both hard drives and machines, um things are going to fail, guaranteed. And they're going to fail a lot, right? You know, if you're running your machine at home, like, okay, maybe I don't know how many people have actually had hardware failure, probably not a lot of people on your desktop machine. But if we, you know, figured out what the odds are and then we multiplied it by thousands, you're going to get multiple failures a week, you know, sometimes a day. And you have to account for that. You know, the first loop the first uh platform we built, uh you know, we'd send it off, we'd say, "Okay, here's all your machines, assume it's a perfect world, go run." And you know, Friday afternoon we'd kick off a training run, go home, have a nice weekend, come back on Monday and find out that um you know, a network card failed or a drive or a hard drive failed Friday evening at 7:00 p.m., corrupted the data or hung the entire process and no work had gotten done the entire weekend. So, you know, if you're doing small scale stuff, like not a big deal, okay, bummer, you restart it, you know, you you go walk away for a few hours. But when you're, you know, again, wasting or using all this money on all this, you know, fantastic hardware 24/7, um those gaps are incredibly important because you lose a lot of time. Especially because it corrupted the run and then we had to go back to a checkpoint and, you know, literally kick the machine and start it again. So, you know, this wasn't rocket science, but it was, you know, work we had to do, which was build in resiliency into the the the code monitoring that training loop as well. So, first thing we did was just, you know, watch to see if something went wrong, you know, if we got a corrupted checkpoint or if if, you know, we were something failed and then we could automatically go back, you know, roll back to the most recent checkpoint, load it up and start the thing again. So, you may lose some time, you know, over the weekend when it fails at Friday at 7:00, but you know, at least it restarted. You didn't have to keep checking on it every moment, you know, from wherever you were. Second thing we did was we actually made the the the training infrastructure much smarter about detecting not only when the whole process went wrong, but when individual machines went wrong. So, we, you know, we looked at what the most common failures were, we, you know, wrote detection code for those and got it to the point where actually could indicate it could discover individual machines or individual, you know, pieces of the data that was getting corrupted, kick them out of the training run and keep it going without any loss of work at all. Because even with that restarting, you know, we still found it was rolling back and restarting many, many times, you know, a day, losing a lot of time just for having to roll back to checkpoints. Okay, so, lots of innovations here, lots of hard work. Monitor's still down, by the way. I think I'm doing fine, but just so you know. Um and uh and you know, I'm revealing all sorts of secrets here that I didn't mean to, but um anyway, we did all this work. It was, you know, this was months and in some cases years of work that we did, but but it paid off, you know, building and evolving this infrastructure. So, if we look at our our profiler again, you see at the top, that's what I showed you before. Um bottom is where it got to. I don't remember when this was from, maybe a year later, that kind of thing. But look at that beautiful chaotic LSD rainbow going on down there, right? That GPU is just running almost full tilt. It's probably running about 80% there, which is actually really the maximum you want it to run at. Going above that is hard to do and, you know, uh queuing theory tell you there's all sorts of reasons why that's a bad idea. So, that's running at essentially as fast as we want it to. Um and, you know, you see down below that uh that uh manager GPU that's doing all the work bringing it together. Um and we've actually got some optimizations there where we said, "Okay, we also figured out how to make the fleet of training GPUs continue to run ahead of the manager GPU." So, very little time lost. Um at this point we had, you know, a training platform that we were very, very proud of. Um and we could, you know, go to the races and and start to really invest in just being the bleeding edge of of, you know, the models and what we're developing the capabilities we wanted. And this worked, right? We've just been 3 years now, we've been doing this for a while. Um the industry has not stood still, right? Nvidia has released, you know, the H100, H200, now the the B series, um you know, like amazing advancements in the hardware, which is wonderful and we, you know, we love to switch and take advantage of those, which just speeds everything up at the bottom. Um the problem is, as fast as our friends have been able to release that hardware, the software side is going faster, right? The industry is is uh voracious. The the the bitter lesson tells us if we just dump more data on it, dump more compute on it, we'll be able to get even more value out of it. Um and so, we got something like this, right? That orange line is is the hardware getting faster, the blue line is is um is really a it's in this case, I think, it's a function of um of the size of the models, right? They are getting bigger and bigger and bigger, which has huge effects on the training, huge effects on the cost and rather than saying, "Okay, wait, great, we did all this optimization, now we can just leverage it and, you know, and run." We have had to, you know, have a significant invest in in a team that um continues to evolve our platform, continues to evolve the capabilities, take advantage of new hardware, take advantage of new libraries or new techniques, change the architecture of our model to try and just keep pulling that line down so that we can get as close to the, you know, theoretical maximum of the hardware so that we're not, you know, sitting around waiting for, you know, machines to finish up so we can take another piece and run it. So, it's been good work, we've done great work. And like I said, um you know, we have uh got a lot of value out of it. Um but it is an expensive proposition. And you know, that is kind of the lesson um kind of instilling you here is going back to what I said earlier, like there's lots of value in doing this if you see an opportunity to build something that is differentiated from the market, that creates opportunity for you, value for our customers, but it comes with a cost and it's an ongoing cost that, honestly, right now is just getting bigger. So, um make sure you know what you're getting into, you know, be willing to learn the hard stuff, which everyone in this room, I'm sure, is. Uh and and, you know, if you if you can do it with the off-the-shelf stuff, um absolutely do it. And by the way, training is only half the process, right? A lot of our GPUs and our money for GPUs is spent on inference as well. In fact, I think it's the bulk of it at this point. I actually don't remember where it is, but um so, it's just as important you optimize that, too. And in this case, there's actually dual benefits. Number one is, yes, you know, the more the more efficient our inference platform is, um the more, you know, bang for our buck we get. But also, of course, if our inference inference platform is faster, um that benefits our users, right? They get the results faster, they're able to move faster, and that's just more value for them, which, you know, ends up driving our business even farther. >> [snorts] >> So, um we invested heavily here, too. I'll talk a little bit about it. Um you know, real roughly what it comes down to is, when you do a training run, especially these large training runs, you build a very big, very sophisticated and complex model. Um you, you know, this thing is big and huge and wonderful, but slow as a result. And so, we go through various steps to try and accelerate it. First, we go through quantization, I always say it wrong, quantization, where we basically look at the data types in there, which are these, you know, very deep, thick floating point numbers, and we say, "Do we need all of the precision in there?" Because if we can get those smaller, um you know, they run a lot faster on the GPU. Now, getting them smaller means loss of precision, which means loss of quality, but maybe that's okay. Right? It all depends on your use case and what you want it for. So, this is a uh task of, you know, running experiments, doing ablations, and figuring out if we switch to this data type, what's the impact on these use cases? And just kind of narrowing in for you and your use case, what's the right balance between size, performance, and quality. Um sparse attention. Uh you know, these models are all built on the core transformer architecture. The core innovation there is the attention head, which basically says, in an image situation, um every pixel could possibly be related to every other pixel in the image because, you know, obviously the eyes of the cat should be the same color, for example. Um so, very expensive part of the architecture architecture says, connect every pixel to every other pixel so they can all pay attention and, you know, together come to an agreement of what the whole thing should look like. It's fantastic. [laughter] It is the core of all the AI that we've, you know, has come into our lives, but like I said, it's expensive. Now, the reality is for a specific use case, you may not need all that connectivity. The upper left pixel may not really care what the lower right pixel is. And so, this is another process of experimentation and finding for our use cases, for our data types, do we need all of that attention? Can we find the ones that we don't need and start splitting them up, you know, slicing them, zeroing them out? And every one of those you do, um that's a savings of performance for the for the GPU. It's able to run faster. Lastly, compilation. The model that comes out, it's a generic, you know, it's a very flexible, versatile model. Can run in, you know, most GPUs out there with enough RAM, etc. Um but we knew we were running on very specific Nvidia GPUs. And so, we, you know, every GPU has its own, you know, optimizations and idiosyncrasies. And we And if you can, you know, make that decision, you can actually optimize and compile your the the architecture and the structure of your uh of your model down to that specific GPU. Nvidia makes a great tool to do this. We worked with them. And again, you know, optimize the thing to run on the GPUs we were running on. Last thing I'll say is don't forget bog standard traditional optimization, right? We have our inference platform over there, which is a bunch of code that is figuring out, you know, how to scale our GPUs and, you know, how to allocate um all the things you do when you're building services and thinking about, you know, dividing it up into services and independent scaling, etc. There's lots of optimization that you can do there. Okay. So, lots of great work there. Hopefully, that gives you a a sense of, you know, the kinds of the breadth and scale of things you have to do and where you need to focus your time. Um we've been doing this for 3 years now. Um as I said before, the ecosystem has been evolving since then. When we started this 3 years ago, everything I just showed you didn't exist in the libraries out there. There were all sorts of libraries to do this. Um you know, the ones that were out in the public were were, you know, designed for small-scale training, really. And all of the big training that was happening, you know, in big companies behind walls, like that stuff wasn't in the library. So, we had to build it ourselves. >> [snorts] >> But time has marched on. All of those tools and technologies have advanced. Nvidia continues to up their game with their tools and their hardware. Other open-source libraries are out there that have taken in, you know, a lot of the optimizations we did. They're all rolling back into the libraries. So, if you were starting today, as I said, like you may not have to do 80% of the things I just showed you. Right? They may be in those tools already. And And, you know, that's like just as we did first, like grab off the shelf, try it, see what you can do. >> [snorts] >> But again, if you're going to do something that is differentiated from what's in the market, if you feel the need for something that is not an off-the-shelf model, there's a good chance that these libraries don't do what you need. And that you're going to find your own bottlenecks. You're going to find your own opportunities to optimize. So, you know, [snorts] as always, like take the gift when somebody gives you. Start with these libraries and then find those places that in order to get the value you want, you can optimize and go further. Um I think that's what I just said. Okay. So, um I think the actually the the main point of this is is interesting, which is we invested all this time and money. We built this model that we were very excited about and worked and and we put it in the market and it it delivered on what we wanted and that was great. But there was another fringe benefit of of taking the plunge and doing this crazy thing as well, which was not only had we built a great model for our users, we had built expertise and domain knowledge and a training platform and an inference platform that really worked. Okay, thank you. Um and that was actually something that turns out we could put to use. So, let me dive into our business in the 5 minutes I have and try and explain that to you. This is our stack, you know, that we came out with. We've got these models we built on the left. We brought third-party models in as well. We surfaced them up into our tools, built on our core, you know, company infrastructure. Great. This is providing a lot of value to our customers. Wonderful. But the more we talked to customers and went out there, we were talking to our enterprise customers and we kept hearing the same question over and over again, which was that's great, but it doesn't quite understand my brand. It doesn't quite understand my, you know, my cartoon characters that that we need to generate. Um we need to go deploy in, you know, some specific regions deep, you know, in parts of the world and it doesn't really understand the culture there. You know, these guys, like they want to tell a cohesive story, especially these big companies, that follow the brand, shows their product, their universe well. And to do that, just as we found, the off-the-shelf models today still really aren't quite good enough for them. Right? And [snorts] so, they need to go through a training process, you know, not as big, a fine-tuning process and with maybe thousands of images, hundreds of thousands of images. They need some GPUs. They need the platform, the people to do it. None of our customers are going to do Well, maybe, you know, 2% of our customers are going to do that. There's some very sophisticated companies out there doing the same thing. All the others basically had no way to do this. But we did. Right? We had sunk costs in building these platforms, these infrastructures, building out this amazing workforce that knew how to use them. We realized we could essentially offer this as a product to our customers as well. And so, that's what we did. Right? We said, in addition to, you know, the core value we provide to our users, we now offer enterprise customization. And we go to, you know, big company, small companies, you know, big brands, and we actually develop models for them. >> [sighs] >> Um same process we said, you know, we go to them, we understand what their needs are, and we ingest all their data, we process it, we enrich it. We figure out, you know, is the pipeline that they need uh just a altered, you know, model dropped into the ones we already have, or do we need to build new pipelines for them? We go through the training process, adapt our, you know, frontier models to understand their specific concepts. And what comes out the other end is a model that they can call all their own and they can use to, you know, accelerate their workflows, etc. Do all the great things AI can do. We call this uh Firefly Foundry. Lots of marketing stuff on here. Please go read about it, you know, especially if you're a large company. Um you know, this has been very successful and it's been wonderful for us to see that we get all these extra fringe benefits from this investment we made. Um you know, not only in offering a differentiated product to the market, but differentiating capabilities as well. >> [snorts] >> We work with a lot of enterprises um on this. We've got a lot of companies. Um a lot of them are big media and entertainment companies. I mostly can't talk about who they are, but um you know, we've had some We're having and have had very fruitful partnerships. I think we've said publicly that we work with Paramount, which has been a great partnership. We're working with Home Depot. Uh working with Disney. So, you know, wide range of different types of companies there. And then all these other ones that I can tell you are really exciting if you could see behind the labels. So, we fold that back into our platform. And now we've got this great end-to-end platform built on our Nvidia GPUs, which for us are running on Amazon, serving up our models, other people's models, and then the custom models. All really unlocked by, you know, that really stupid decision we made 3 years ago, which is we're going to build this from scratch. And so, if you're here listening to this talk, I think maybe you're wondering should we make that same You know what? Okay. I got a minute. I'm going to show you some quick examples before I jump in here. Um this is an example of something we did with a customer. They wanted to be able to take uh you know, a back painting and then a a sketch a sketch artist did of their their characters, right? Their IP, and be able to give a prompt and just generate a video for it. And so, we did. And if you saw their original stuff, like there's lots of ways you can interpret that sketch. This looks exactly like their IP to the point where they could put this out actually in the market in some places as as real content. Another example here, this that one was characters. This one is look and feel in the style. Um given the sketch, given the starting point, and say, okay, I want this thing to bend upward. It does that. Two more bits of eye candy. This is uh some open-source IP that I think Blender and Netflix put out. Uh we did a custom model on this, learned all these characters, the look and feel. This is all actual source material that you can find online. Um and then we generated a bunch of uh videos that all are, you know, you can imagine are great content that you know, this is all relatively low-res stuff because it was a proof of concept, but, you know, the quality of what I can show here is absolutely within the range of what many of our customers would say, "Awesome. That's, you know, final pixels for some use cases." Okay. So, what I was saying before, the end result here, happily ever after, at least until the next big, you know, innovation comes along. Um if you're here asking should we do this, remember those basic questions, which is, you know, there is huge value you can get out of doing custom training and building your own models. Um and uh you should absolutely take it on if you believe there's the ROI there. In order to do that, understand what it means to diverge from the from the stack libraries. Use those. And then if you do decide to dive in, have fun because this is really fun stuff. So, I think that's it. Thank you guys.

Original Description

Learn how to design, train, and deploy large-scale custom generative AI models. This session will step you through data preparation, fine-tuning, and infrastructure management, and will highlight best practices for scaling large models efficiently. You’ll also see real-world examples with Adobe Firefly Foundry and NVIDIA NeMo integration that demonstrate how teams are applying these techniques to production-grade AI systems. Speaker: Ely Greenfield | CTO | Adobe Key Takeaways: Data Preparation: Discover how customer images, video, and 3D assets are analyzed using vision language models to extract unique world context and specific visual style. Fine-Tuning: Learn how very large generative AI models are fine-tuned at scale. enabling distributed training across tens of thousands of customer assets, using NVIDIA's technology stack (including CUDA and NCCL). Infrastructure Management: Understand how AI infrastructure is built to serve large and varied customer bases, each with dozens of personalized models deployed concurrently, while maintaining secure isolation, performance efficiency, and production reliability. Industry: Media & Entertainment Topic: Agentic AI / Generative AI - Video Generation Technical Level: Technical - Intermediate Intended Audience: Developer / Engineer NVIDIA Technology: Cloud / Data Center GPU, NeMo, NVIDIA AI Enterprise #NVIDIAGTC
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
Unlocking Open-Source AI: 5 Tools for Unbeatable Privacy and Cost Efficiency
Unlock 5 open-source AI tools for enhanced privacy and cost efficiency in tech companies
Medium · LLM
📰
The AEO tricks that don’t work in AI search
Learn which AEO tricks are ineffective in AI search and why, to improve your search strategies
Medium · AI
📰
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics
Learn to build production-grade LLM evaluation pipelines to catch hallucinations before deployment, replacing manual 'vibe checks' with automated metrics
Dev.to AI
📰
Building AI Data Pipelines — How to Feed Your LLM Fresh Web Data
Learn to build automated AI data pipelines to feed your LLM with fresh web data, saving development time and reducing maintenance headaches
Dev.to AI
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →