High-Efficiency Diffusion Models for On-Device Image Generation and Editing [Hung Bui] - 753

TWIML AI Podcast · Advanced ·🎨 Image & Video AI ·8mo ago

Key Takeaways

Hung Bui discusses high-efficiency diffusion models for on-device image generation and editing, including SwiftBrush and SwiftEdit, which enable high-quality text-to-image generation and editing in a single inference step.

Full Transcript

So we realizing this open weight models 7 billion parameter model we're still getting complaint from the community in Vietnam that oh this model is too big right we can't fit in in on on our GPU right and we said okay fine you know let let us go you know one more step try to like reduce the number of parameters try to half the size the model to uh less than four billion parameter and with a couple of improvement over the way we train it we noticed that this model less than 4 billion parameters actually perform you know even better than 7 billion parameter model. All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Cherington. Today, I'm joined by Hung Buie. Hung recently joined Qualcomm as VP of technology through the recent acquisition of Vinai Research, which ranked in the world's top 25 industrial AI labs based on research output in top conferences like ICML and Nurips. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Hung, welcome to the podcast. >> Thank you so much, Sam, and it's my great pleasure to be here. So, we've got a bunch of really interesting topics to dig into, including your research into topics like diffusion models, image generation and editing, and more. And of course, how to make all of that efficient on mobile devices. To get us started, I'd love to have you share a little bit about your background, which includes time at places like Google DeepMind and Adobe Research, among others. And tell us a little bit about how you got into AI. >> Actually, let's see. I I did my PhD almost 30 years ago and the PhD is actually on the topic of multiation system and back then uh it was also very interesting AI topic. Uh and the the reason I got into AI is is is just by curiosity. During my undergradues in Australia, I I learned about things like touring test and I was, you know, very curious to myself that, you know, how be able to program the machine to pass the touring test and and back then I have to be honest that I don't think that I'm going to uh live to see the machine passing the training test which, you know, [laughter] it's kind of like taking for granted now. tell us a little bit more about some of the work that you've done in your career. >> So, um I think my first job uh first real job after academia um was at a place called the AI center at SRRI International Sri Stanford Research Institute. >> That's right. Yeah. It it is formerly known as Stanford Research Institute. And back then, you know, we're talking about 20 years ago, I got a chance to work on some project called Kalo. And back then we try to build, you know, guess what? A personal assistant. It was running on a desktop. Uh, and we were trying a lot of things. Uh, you know, a lot of machine learning techniques, a lot of proistic reasoning techniques to understand the intent of the user and so on. um you know we even have a system that can actually record all the user actions uh on the desktop screens and you know uh it was it's I think a very interesting project uh I think back then it was sort of like the largest AI project at the day and yeah I think one of the you know artifact of that project was a little system known as Siri right and >> right >> and after that I I uh start to move to various industrial research labs in the Bay Area. I spent some time at Nuance lab on natural language understanding. Uh also some time in Adobe research where we start to look into applications of uh machine learning in various areas of the business and and this is this is I think this is time that charity AI starts to become popular with models like you know variation autoenccoder and gun and so on. Um and it's also correlate with my next move to uh DMI. Um and I was part of the DMI team also still based in Mountain View the Bay Area. And I think there was that was really interesting time when looking back uh but yeah I think uh you know in 2019 um I actually uh got an opportunity to you know move back to Vietnam which is you know the place I was born and had a chance to set up you know the first AI research lab in the country and >> tell us more about that like you you're coming from some of the top industrial research labs in the the Bay Area you know, the place where it's all happening, like what prompts you at that point to to leave all that and go back home? >> I mean, I I I could tell you that, you know, it was a, you know, very exciting opportunity. I you know, I think a lot of us, right, um, when we're sort of like, you know, working for, you know, like big company in US or so, uh, I think we we all have a little dream, right, you know, to return to the whole country and and and make an impact. And I think that that's kind of like a big opportunity presenting itself for me. But I think you know in hindsight I I I I should admit that I also you know took a big risk as well because it wasn't known whether it's actually possible to actually run a a you know like a a proper worldass AI research lab in you know place like Vietnam right because it is known that that you know AI research is going it's like it it's something that that requires huge investment especially around talents so yeah um big risk but you I'm I'm glad that uh that I did it and uh you know it paid off. So, >> and how did you craft a research focus or a direction for the lab? I must say that uh you know I I must credit this to the really interesting research that I was doing uh back in the days where I was a deep mind right so uh I think it is quite natural that you know uh we continue to follow the same direction in what is called back then deep generative models which is kind of like the uh precursor of generative AI today. So you know we you know from day one at Vaii we start off working on things like you know variational autoenccoder generative adversarial network and of course you know like auto reggressive models model for Vietnamese for example B model for tweets and and those are kind of like you know just the natural research topics that people latch on back in those days. uh it it just turned out to be that all of those topics you know are super important and it really prepared us really well uh for the you know uh the evolution of generative AI that follows until today. >> You mentioned that talent was a big challenge and starting a lab there in Vietnam. How did you address that? I I remember myself that uh during one of the first few weeks when I was here in Vietnam, I stood in the Hanoi office, I looked out of the window and uh you know, I kind of like you know just checking the view of the streets of Hanoi and and that that's when it hit me that okay, I'm no longer in the Bay Area, right? [laughter] And and and then Right. Okay. So, how do you you know, how you actually build a team here? Uh and um but I I think you know one thing that I know for sure is that um uh okay I have to be able to convince people to move to Vietnam and work in Vietnam, right? for for for for for for this effort to be successful. And then I have to be able to have a a good balance between the experienced guides, right, who kind of like, you know, been there, done that, but still willing to move back to Vietnam uh with the u I would say the the young talents, the people in the country, right? the talents are still in the country whether you know it's just really really smart young guys might not be that experienced in in AI or you know things like deep generative models but you know smart enough so that you can actually train them quickly so you know in incredibly that's that is that was the the strategy that you know we went about in building the team right so we were be able to hire a couple of uh you know I think really strong uh research scientists with you many years of experience working in the US and also other countries. they willing to you know um you know move back to Vietnam and and work with me to build up the team and then for attracting the young talents then we we actually started uh the first AI residency program not only in Vietnam in Southeast Asia actually and yeah and and that you know has been a a great uh way for us to attract and recruit the best young talents uh in this part of the world and you know they they come and work with us you know become a full-time employee with the company for uh two full years and yeah uh you know we we've had I think uh a really fantastic opportunities to work together with those young talents they they just you know they're so smart and they're learning things so quickly >> so the lab eventually became known for its work on efficiency um So uh mobile device getting models to work on mobile devices and that ultimately led to the acquisition by Qualcomm. I think that's you know a clear shared interest you know from the your your initial research focus on deep generative models. How did this focus on efficiency come about? >> Going back to this is 2019 right? Uh this is al also the time when people start to look into how to scale up these models to work with larger and larger quantity of data and how to get model of you know larger and larger uh size right the number of parameters keep on increasing and and demonstrate that the bigger the model like the the more capable it is right and you know in Vietnam I think we try to follow the trend right but very quickly uh we are limited by our access to computational resources. we we have you know the the investment uh to to to have access to a small TPU cluster uh locally in Vietnam right but you know very quickly you know we we cannot compete with the giants in in big tech um and and and because of this right so knowing that technology like deep generative model is important or generative AI is important uh but yet you know we kind of like you know we know that uh we are being constrained by the access to to computation resources to to to do to train these models. Um, and that that led us down to the the only path that you could possibly go, which is, you know, how you be able to make these models more efficient, which means that we have to figure out clever way to make the smaller models work, right? Get more chuso out of the smaller models. And yeah and and I think that that's the reason why efficiency is a very natural focus. Right. >> Talk a little bit about some of the research that that path led you to. Was it primarily focused on topics like quantiza quantization and and related ideas or you know how did you uh how did you approach it? I think the first thing that we're looking at is what we can actually get by smaller models, right? And and let let me give me let me give you an example. Chip was released. This is roughly November 2022. And this, you know, obviously caught all of us by surprise. Uh but we were looking we we were already working on model like BR for Vietnamese, right? and and and then we're it's very natural for us to figuring out okay so can we actually pre-chain something like uh chacht but using only Vietnamese data and we know that this is at a time where chacht I think ch3 is close to 200 billion parameter right and we know that there's no way be able to to to ri that that scale in terms of size uh you know with with the training resources that we have so we kind of like kind of perform an experiment to to see what we can actually do, right? With a very small model, something only 7 billion parameter back then, right? And you know this is actually going against the trend when you know other companies was trying to get you know larger and larger model. We we kind of like you know okay we ask ourselves okay what is it you can do with like a seven billion parameter model pre-chain completely from scratch right using Vietnamese only data right so this is kind of like you know first it's first it started as an experiment >> was this based on the GPT2 architecture or what what was the model architecture for it did you come up with that uh independently as well >> the model itself is is is well understood It it's just like we're not entirely sure you know seven billion parameter and the Vietnamese only data like whether we can actually produce you know whether we can actually produce anything interesting right and and so we you know obviously we grab all the Vietnamese data that we could from the internet uh crawl right and and then we you know fit into the model architecture and I think it it took us you know a couple of months, right, using our compute. And the end result was something that actually surprised us because we thought, oh, you know, we actually have a model that can speak Vietnamese really well. You know, you can answer questions in Vietnamese, you can, you know, write letters in Vietnamese, it can, you know, like, uh, write poems and and and all of that, right? So, so things that that you see with the earlier version of CHPT, we actually saw it in Vietnamese with only a seven billion parameter model. Um then the next thing that we did is that we asked the question what if we reduce the number of parameters for this model even more right so it's kind of like a counterintuitive back then right it's it's so we asking oursel the questions like can we get more for less right here just less number of parameters or can can we get the same same performance with less number of parameters u like all of this is because you know it's it's is the it's the focus on efficiency less number of parameters I I remember that that that back then with even with 7 billion parameter model right the other people in Vietnam right the other team in Vietnam they still complaining right so we releasing this open weight models 7 billion parameter model we're still getting complaint from the community in Vietnam that oh this model is too big right we can't fit in in on on our GPU right and we said okay fine you Let let let let let let let let let let let let let let let let let letlet let let let let us go you know one more step try to like reduce the number of parameters try to half the size the model to less than four billion parameter and with a couple of improvement over the way we train it we notice that this this model less than four billion parameters actually perform you know even better than 7 billion parameter model. So we already know that there are uh you know a lot of room to optimize the model the way we train it and getting more uh more performance uh out of model that is even smaller. >> Can you talk a little bit about some of those optimizations and how you tweak that training recipe in order to get decent performance out of the smaller models? So one thing is is um how do we get even more data right now we're not going to get more data obviously right but we can iterate through the same data multiple times and back then we already thinking okay you know can you actually use synthetic data to even try more pre-training right >> you found everything you could find [laughter] >> okay so later on of course you know like uh you know other people found the the same thing for for the internet but then back then because we limit ourselves to Vietnamese and we we already hit that that boundary. Okay. So we figure out okay you if you iterate over the same data set you know more right you know perplexity keeps going down and and the model and the model keeps getting better right and and also there there are a couple of minor adjustments on the model optimization side that we did but but yeah I think at the end we have a sub4 billion parameter model and it starts to fit into you know just the the basic GPUs that uh you know people in the lab speak people in universities in Vietnam they could use, right? And and so that that was a very welcoming um addition that that we support the comm the local community here by, you know, just just giving this this models uh giving them out as open weights. >> I'm just thinking about how so much of the innovation in models is driven by data sets and here you are trying to create uh models that perform well on Vietnamese language. you collected kind of the um you know just the the raw corpus of uh internet information in the Vietnamese language but um you know when I think about you know traditional kind of NLP and you know model research there's so many benchmarks that folks are trying to optimize models against that teach the models you know new and different things and you didn't have all that for Vietnamese. Did you create any of that uh or did you find that just you know kind of this the raw training on the the data that you scraped was enough? we don't have data for Vietnamese for evaluation and so on right um so we we had to to create some of that ourselves of course right uh but I think lucky for us uh a lot of the corpus for English can be auto translated into Vietnamese right and and I think looking back machine translation between English and Vietnamese and in particular we actually also were work also also working on machine translation between English and Vietnamese, right? That was working well enough. >> So that let you bootstrap take an English QA data set and translate that into >> just Yeah. >> Yeah. Exactly. Exactly. Yeah. >> And so how did you how did you get into the image generation side of things? >> So text and languages isn't the only focus in the lab. Right. At the beginning we got people working on for example GAN model for image generation. Then of course you know the thing that comes around is uh diffusion based approach. the quality of image generation that you know methods like diffusion model can actually produce you know at the time was getting more and more impressive and and so of course you know uh because we were working on image generation using GAN a lot of people in the lab starts to you know experimenting with the the noising diffusion and that's that you know it just you know it's it's a very natural thing for us regarding efficiency it's a little bit different right I mean the goal is the right? How you be able to you know get text generation and also here's how you'll be able to generate an image but in a most efficient way. Um so so interesting enough for image generation the size of the model isn't that big right compared to to sort of like you know large language model but for image generation especially a den noising diffusion approach the main thing was that you need to run this den noising steps for many many times right so it's you know it's a sizable network but you have to repeat it for many times you know sometimes between 50 or 100 time steps. What that means is that uh you know if if we have a you know a big compute you still have to wait right back then I think if you run this you know uh with chat GBT you still have to wait uh it's it's not real time um you can you know this is noticeable lag uh and and if you have already small compute cluster you have to wait even more right so so for us to be able to do experiments then we have to look at ways to to to just be able to generate image in a much shorter time step and and and significant reduced latency. And and that drives us down the path of asking ourselves the question, can you actually get a model that can share an image with as few number of steps as possible or even with in just one step, right? So, so imagine that if you can just do this, right? in just one step uh your latency is you know all of a sudden boost up by almost two order magnitude. >> Is it still diffusion at that point or are you needing to create an entirely new approach? We are working on the assumption that you already have access to uh the noising diffusion model, right? So you already have access to this model that can actually produce beautiful images >> but it'll process. >> Yeah, exactly. So so you assume that you have that and and you would treat that as a teacher, >> right? And you use kind of like a distillation approach, right? Maybe I I I I should uh you know like explain a little bit in about intuition here. So distillation is usually thought of as you know you you you have a teacher that is a big model and you and this big model encapsulate a lot of the knowledge about the task at hand and you want to distill that knowledge into a smaller student. Right? But and I think you in here when you when here the problem is kind of like a step distillation because you don't have a big model. It's just it's model of the same size. You need to run it for multiple steps, right? And and here you kind of like have to you know uh uh distill that multi-step knowledge right into something that that has just one step. >> Meaning because your diffusion model is already relatively compact. It's it's not a huge model. You're not in this distillation process necessarily trying to shrink it. You're just trying to uh eliminate the steps. >> Exactly. Yeah. Yeah. In fact, the architecture of the student and the teacher, you know, it could be almost the same, right? It's just that the teacher you have to run multi step just one step and and so another way you can you can kind of like think about it is that so first of all the teacher is given right. So this is this is already a really strong uh den noising diffusion model that has already been trained with lots of data right and it can actually generate beautiful images with you know like let's say 100 steps and and and it is given right but the way it is given to us is through a den noising function right because you're going to have to repeat this den noising function like a 100 times right so so that's fixed right that that network is fixed right we we're not going to change that Okay, we would have to learn a student network, right? Okay, so that's distillation a standard distillation, right? For any student network, right? You can also estimate a den noising network for that student, right? This is a function of the student network, right? And now you're you are forcing this denoising function for the student network to be the same as the denoising function for the teacher network. By forcing me the same, I mean are we going to minimize loss >> constraint? You're going to minimize the loss to get them to be as close as possible. >> Exactly. So, so yeah, the teacher network is is fixed and constant, right? But this secondary teacher network or the coach network needs to be learned. And what's the intuition for why you need the secondary coach network? Like if the teacher has all of the knowledge about how to generate the images, what is the secondary teacher actually doing? >> Great, great, very good question. Uh very good question. So, so this this the answer could actually get to the heart of why you need that um additional instrument, right? So uh remember that that in the first approach right we are we we are asking that the den noising function uh for the teacher right has to work really well for on on on on distributions that generated by the student which would be true if the student would generate the same distribution as the teacher would. So this is would be true if if if we we manage to successfully learn the the distribution generated by the by by the teacher. So so we we we I think that there's like agreement that yeah okay this thing should be zero you know at convergence. Okay. But during the beginning uh of the process, right, the distribution generated by student is widely different from the distribution generated by the teacher, right? >> Okay. >> And and and that and and so not very good yet. >> Yeah. Yeah. Not very good at right. So so basically the signal from the den noising function from the teacher uh might not be a very strong even a good signal to follow. Right. and and and that is the intuition, right? >> So the secondary network is kind of acting as a bridge between the students the early student distributions and the teacher distribution. >> Yeah. Yeah. Yeah. So that's exactly right. So early on in the process, right, uh we would want to explicitly estimate the denoising signal for the student, right? Um and then we minimize the difference between the denoising function versus the the deninoising function of of the t-shirt and and that's that's kind of like providing the bridge of of uh guiding the process um uh during the early stage of the optimization process. And so that initial work again this is swift brush and that is focused on um you know we've got image generation diffusion it works great takes a long time you know 100 steps so that's a lot of kind of inference so how do we make that more efficient well let's skip the hundred and do like one shot from noise to the to the desired output image um you know even with all of the explanation of how that works in the intuition like it's hard to believe that it actually works and works well you know talk a little bit about qualitatively what kind of results you saw >> again you know we we we're getting very good quality in terms of image image quality right um and and also quantitatively as well right with all the benchmark that we measure especially with you know some additional improvement that uh you know we later on did for the second version of sheep brush we call it's just simply you know a slightly improved version of sheep brush um the the scores that we getting is almost as good as a score sometimes even better than the score of the original teacher >> and which benchmarks are we talking about >> it's it's a bunch of benchmark it's very standard benchmark on image quality um and also diversity right um and and and yeah at you know um and and and we this is kind of like a standard benchmark that you'll see in in any papers in this area. Um and uh you look at the score for uh this oneep model you know we notice that you know we're getting the score you know sometimes as good as as as as a teacher itself. Uh the challenge is to get this thing to converge, right? The challenge is to be able to get this training pipeline uh to to to be in in in a stable condition and you know on the initialization and also the training pipeline so it actually converge but but but once it converge you know it's uh uh we noticed that that results uh are really strong. Yeah. So, so then you know like okay um you can ask the question that okay now that you can do text to gener text to image generation in just one step it is really efficient and so on right uh can you also do text editing in a similar uh efficient manner right and you know that gets such into this topic of uh image editing which is you know I think it's is it's it's a it's a topic of of a recent paper that we just published at CVPR this year >> that was a swift edit paper Yeah. So, S brush which is just you know generating the image but then SIF edit right is you know be able to also edit uh images uh quickly as well in in just like one step right so that that's the key is the SI repep image generation swift edit uh one step image editing that is the goal so you know um you me to get into like how how it work in kind of like a high level >> yeah assuming that you already have a very efficient one-step text to image model. Getting to text editing, you know, it it's a lot more simple, right? If you don't have that, then you know, forget it, right? [laughter] But but if you already have the onestep text to text to image model, then you know text editing, you know, you can kind of like start to to to to conceptualize things in in in a much more uh simpler fashion. So, so, so the the the way the way we we architect this this model is is as follow, right? So, you have uh an original image, right? Of course, you have a text ROM and you you want this text from to kind of like, you know, operate on the image, right? And get into the the final image here. Okay. But the way we want to do it is that first let's get this text from into noise, right? And then we already know how to get from noise to the to to the image, right? Given a text, right? So so so so then then then the challenge here is then how to get from image to noise, right? Condition on a text. Okay. Yeah. So to get from the image to noise, right? You could kind of like okay, you know, we could set up various laws to do this, right? If if you have real images, right, as training data, right? you can set up what is called an inverted network or inverted model right and again this is a one-step model. So hey you know this is like a neuron network architecture that would take you from image right to of course the encoder in that image latenc the the representation that image to noise but you should do it in such a way that if you apply sip brush again remember we have access to sip brush right if you apply sweep brush again it would take you from no to a latent representation on the same image, right? So, you know, like latent representation of that image Z of the noise, apply C brush and you get another latent representation and and you can simply say that okay, you know, this lat these two latent representation has to be the same, right? And and and that that gives you a very natural loss. You you can also go to uh the image itself, right? You can go like okay from this this this Z latent representation you can use the decoder of course you know all all of this coming from from from S brush right to generate the actual image and you can say that okay the the the the the original image here and the image that you generate by applying Sbrush to to this intermediate noise right has to be the same right and and and this is this is this is if you have access to to to to only uh real training data, right? But you can actually do more, right? Because you you you have sweep price, right? So you can you can syn synthesize any images you want. So so so you can do this without access to any real data as well, right? So you can go from noise applying brush going to a particular uh image or the latent representation of the image, right? And from that apply this inverted network going back to noise, right? And you can say that okay this this the noise is starting from and the noise that that you ended up with as a result of applying the inverted work has to be the same right so this two epsilon you know and epsilon hat here has to be the same. So that's yet another loss function that you can that that you can have right as another signal for training this this inverted network. So having access to an efficient onestep model right uh allows us to kind of like in train this inverted right work to invert from latent resetting Z to noise and and and and and and and do that quite efficient right because you can differentiate through this network quickly right it's just one step so you can differentiate through this network really quickly and you can also combine signal from real data also signal from completely synthetic data because the synthetic they can be generated really quickly with this onestep network. Right? So, so all of a sudden you have combination of really just you know very you know like loss function that does that that does that's that's highly intuitive and and most importantly you know you can implement it and train it very efficient all thanks to you know already have access to this onestep image generation model which is swift brush which is like I think the core to to making this sweep edits uh possible. We'll link to these papers in the show notes and I encourage folks to uh pull them up. In particular, the SwiftEdit uh paper like the the images are super high quality and it's surprising how high quality they are given that process. >> Yeah. Yeah. Yeah. It also works really fast, right? But it um um so we can actually run this on you know a standard um one single GPU and you know it it took us like I think a fraction of a second I think uh a quarter le something like that second yeah to do one uh uh one image editing which is real time right I think even before you finish typing right you already see the resulting image >> so that work is all focused on efficiency I know you're also working on kind of agents and uh making that type of model efficient enough to run on mobile devices. Can you talk a little bit about the way you're thinking about that space? The way I think about this is is that you I think we all want to have uh agents assistants that are personalized, right? And and and what that means is that this agent or assistance they need to have access to our private information otherwise otherwise how can it actually be personalized, right? And where are the private information reside today, right? Well, you can argue that a lot of those private information actually resides on the our personal devices, right? But that also means that privacy suddenly becomes really important. And so the way we think about this is that we want this agents or assistant, right, to be able to do as much of the computer workload on device as possible. Why? Because you know uh first of all it's very close to the personal data that's sitting there and secondly more importantly right if you can actually process all that information on device right and you don't have the risk of you know exposing this private information to third party and so on and of course you know if there are tasks that actually require information you know from the internet right and then these agents can also you know collaborate with you know a bigger agent right a more sophic more sophisticated agents with access to broader knowledge from the internet right that c reside on the cloud right >> and so how does that broad direction or vision translate into specific research projects >> so so first of all this is kind of like you know I want to kind of like say that this relates very closely to all the stuff that we talk about on efficiency right to be able to run this model to support this you know aentic behavior entirely on device right means that You again you have to look for ways to have not only large language models but you know this is large multimodal model models that can actually take in both text and also you know multime media as well. Those are the things that you actually have access on a device today, right? And you want to run all of that on a local device, right? So that means a lot of experimentation with uh smaller models, right? uh models that ranging from four billion parameter or even less right and various way to get these models to perform efficiently in terms of you know the rate and you can actually uh in chest tokens prefield token rate versus the you know encoding also also decoding token rates right um another thing is is is that I think we we start to kind of like look into what are the source of information right uh that are really important for this ondevice agent right so I I I mentioned access to private information on a device right and making sure that you process information securely on the device itself but but then I think uh you know very soon right you would need to uh look at okay the information that you're getting not only from your phone right but from other wearable device, for example, smart glasses, right? And and and and that's that that that's that's very rich in terms of multi modality and and that that just opens up like a really interesting uh space of research, right? And how do you enable this model so they can actually, you know, understanding uh video that's coming through your smart glasses, you can actually understand the content of the screens, right? That's actually uh the user looking at in in the phones and so on. and and yeah to be able to kind of like you know distill all that information to a a compact representation to be able to do that in a way that's efficient so that you know you you kind of like you know don't consume a lot of battery right and then store that information somewhere so that can actually be retrieved right and so all of that is just a lot of work that's you know needs to be worked out both in terms of again you know like the locking information uh compressing this information uh bringing it to a form that can actually be ready to be retrieved later, right? uh and and and for retrieval you can think about this this as okay well yeah you know something like brow but you do brow but not only you know have access to information but also installation on the device right but on the device is being represented in a different way the data is different so you're going to have to make it work for this this this new distribution of data as well >> when I think about agents and the kind of broader innovation that's been happening there Um the rise of reasoning models has really changed uh that game and what's possible. Um you know that's very inference and computeheavy. Um I have to imagine that that poses a big challenge to running those kinds of models locally on the device. You know are you thinking about um you know this idea of inference time scaling and what that's going to mean for ondevice models? Thanks for bringing that up. I mean infant scaling is a really important topic. uh you mentioned that yeah it's a challenge to run this more run this kind of technique in infant time scaling on device but but I the way the way we should think about it is is that uh it's both a challenge and an opportunity right uh why I say is an opportunity is because there has been multiple works in the literature showing that if you take a small model in terms of number of parameter and you apply test scaling Right? And you measure on a particular kind of task, let's just say math, right? And with test scaling, the small models actually, you know, perform a lot better than model that significantly larger, right? And and and so test scaling is a way that you can actually make the small model, right, to even bit a much larger model on a specific task, right? And and I think this this is this is this is really important because all of a sudden it enriched the capability of the small models. >> It's an interesting give and take. So like the test time scaling implies that you're exploding inference and that's a you know a constraint on a mobile device, but it also inherently allows smaller models to match the capabilities of much larger models which is a a tailwind for you trying to get this running on a mobile device. Yeah, exactly. Exactly right. And and of course I mean regarding the the challenge of how to you know make this work efficient uh on a mobile mobile device and and this is this is kind of like more compute b right it's no longer a memory b it's like compute b u this is this is a topic that that that we are looking at very closely right how to again make this you know test time scaling work more efficiently for example how do you do it assuming that you have an upper bout in terms of access to compute Right. Uh it's almost like a fixed compute budget. So you have a fixed compute budget and then then what other strategy >> how do you allocate your resources across? >> Exactly. Exactly. It's interesting because I was going to ask the degree to which test time, you know, test time scaling and getting these reasoning models working on constrained devices is different than just the inference problem of, you know, having any LLM inference happen efficiently. And and so it raises these kind of, you know, meta issues. And and a good example of that is if you've got a fixed overall, you know, budget, you know, whether that's, you know, compute or latency or whatever, and you have you're able to do some kind of planning that this is going to require some number of uh inference requests like how do you optimize where you spend your compute? Uh is an interesting right, >> you know, way to formulate a research question there. >> Yeah. Yeah. I think Yeah, I think you you you said it pretty well, right? and and and and in in in in a sense you know uh this test time compute kind of like combine the probability that uh this is kind of like you know has been learned with an L&M which is kind of like almost like a forward predicting right uh with the kind of like an estimation of what's the future reward is going to be right so it's kind of combining probability and utility uh in a sense so it's kind of like you already gone over what the original L&M was designed to do right so in that way it's I think it's is you know it it it it allows us to you know move into a a much I think a richer problem right and and and which is you know how how we able to to find a particular answer path that maximize expected utility um and and that that's a very rich a rich framework right so you know things like you know compute resources resources constraint you can you know you can see that Yeah, you know it's possible it's plausible that you can actually formulate this uh under the framework of uh you know optimizing for future expected reward. >> And so the this idea of uh agents uh you know clearly kind of opens up you know many different research directions as you know and it kind of serves as maybe a grounding kind of application area. Are there others that are high priorities for you? I I would say that um yeah uh in your in general um efficiency it's important for many broad topics and I think it's this is something that that that people uh also have realized you know like I think for for for some time you know they you know people like probably don't focus too much on this issue efficiency but I think recently I think you know they they you know there's a lot of focus on on this particular topic. So I I probably don't need to say anything more. So it's been I think six months since the acquisition. Um just uh what are your biggest kind of lessons learned through that process and uh what are you most looking forward to as you you look ahead? I I I I must say that um uh we we we we were pretty lucky because um you know issues like efficiency right and and ondevice uh uh models is is something that that is just you know something that Qualcomm and Qualcomm AI research those folks you know u have been working on you know similar problem for you know like quite a while as well and there's a great depth and breadth of expert ities um within Hong Kong research itself. Um so no because of that right we uh you know I think uh integration um um has been um you know a lot smoother right uh um we don't have to you know change the objectives in terms of you know the way that we've been working it's more or less uh the same uh focus uh uh so for us it's I think it's just it is more of a matter of of you know learning think about the capability within Quongcom a research AI research and see how we can actually best help uh and and enhance the capability of the group. So yeah um I think you know we we were very uh fortunate and also you know like you know having access to to um the talents in the resources area research is you know like uh making us even more excited >> and talent was one of the big challenges that you mentioned when you got started. Are you continuing the residency program? Ah yes uh absolutely I think this is yeah so this u um uh the the AI resency program uh uh today uh we uh this something that uh will continue um to reinforce um and you know like a little bit of a of a history of the residency program we started this um uh hiring the first batch in 2019 that's six years ago Um and yeah and and and uh at any point in time we we have about between 40 and 50 residents in the lab. Right. >> 40 and 50. >> Yeah. Between 40 and 50 residents. >> That's a lot bigger than I imagined. >> Right. Yeah. Uh they say >> how many researchers total in the lab? >> The total number of people in the lab is is about 90 people. Right. So, so the number of resident is almost half or even a little bit more than half of the of all the research and and engineers. >> Yeah. Wow. >> Because they stay with us for two years, right? they have enough time to contribute uh significantly to to to to our research and also engineering process and even now right the people who have gone through the residency program uh you know almost like close to a hundred of them and and a lot of them are actually in you know top AI PhD programs in the US or you know Europe, Australia and so on and I think that that that has been a really nice tradition And we, you know, want to continue to keep it that way. The program now would continue under the branding of new branding of Quangom AI residency program. And I think we just hired the first batch of research residents and we we continue to look for ways to improve and expand the program as well. And we are about to recruit the first batch of engineering residents. So, AI engineering residents and and and I think this this will also provide it the opportunities for the young local talents right to have this unique experience to be part of an AI research lab of a big tech company like Hong Kong. >> I don't know that we have a huge listener base in Vietnam, but uh in case there are folks listening that might be interested in the program, is there a a page that they can go to to learn more? Yeah, we have a landing page for the Quong AR resency program. Uh you can find out about it from Quong website itself. Uh and from that uh we have link to recruitment. >> Well, we'll find the landing page and stick it in the show notes. >> All right. Okay. Awesome. >> Awesome. Well, Hong, it's been great catching up with you and hearing a bit about your journey and the the projects that you have worked on and are embarking on as part of Qualcomm. Thank you, Sam. And again, you know, thanks a lot for RTBT to, you know, uh share my uh thoughts and uh opinions here. >> Thank you. >> Yes. Thanks so much.

Original Description

In this episode, Hung Bui, Technology Vice President at Qualcomm, joins us to explore the latest high-efficiency techniques for running generative AI, particularly diffusion models, on-device. We dive deep into the technical challenges of deploying these models, which are powerful but computationally expensive due to their iterative sampling process. Hung details his team's work on SwiftBrush and SwiftEdit, which enable high-quality text-to-image generation and editing in a single inference step. He explains their novel distillation framework, where a multi-step teacher model guides the training of an efficient, single-step student model. We explore the architecture and training, including the use of a secondary 'coach' network that aligns the student's denoising function with the teacher's, allowing the model to bypass the iterative process entirely. Finally, we discuss how these efficiency breakthroughs pave the way for personalized on-device agents and the challenges of running reasoning models with techniques like inference-time scaling under a fixed compute budget. 🗒️ For the full list of resources for this episode, visit the show notes page: https://twimlai.com/go/753. 🔔 Subscribe to our channel for more great content just like this: https://youtube.com/twimlai?sub_confirmation=1 🗣️ CONNECT WITH US! =============================== Subscribe to the TWIML AI Podcast: https://twimlai.com/podcast/twimlai/ Follow us on Twitter: https://twitter.com/twimlai Follow us on LinkedIn: https://www.linkedin.com/company/twimlai/ Join our Slack Community: https://twimlai.com/community/ Subscribe to our newsletter: https://twimlai.com/newsletter/ Want to get in touch? Send us a message: https://twimlai.com/contact/ 📖 CHAPTERS =============================== 00:00 - Introduction 04:44 - VinAI Research 07:03 - Building an AI team 09:02 - First AI residency program 09:47 - Focus on model efficiency 11:49 - Training a Vietnamese LLM 16:29 - Model optimizations for Viet
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from The TWIML AI Podcast with Sam Charrington · The TWIML AI Podcast with Sam Charrington · 0 of 60

← Previous Next →
1 Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
The TWIML AI Podcast with Sam Charrington
2 How to Build Confidence as an ML Developer with Siraj Raval - #2
How to Build Confidence as an ML Developer with Siraj Raval - #2
The TWIML AI Podcast with Sam Charrington
3 Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
The TWIML AI Podcast with Sam Charrington
4 Interactive AI, Plus Improving ML Education with Charles Isbell - #4
Interactive AI, Plus Improving ML Education with Charles Isbell - #4
The TWIML AI Podcast with Sam Charrington
5 Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
The TWIML AI Podcast with Sam Charrington
6 Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
The TWIML AI Podcast with Sam Charrington
7 Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
The TWIML AI Podcast with Sam Charrington
8 Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
The TWIML AI Podcast with Sam Charrington
9 Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
The TWIML AI Podcast with Sam Charrington
10 Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
The TWIML AI Podcast with Sam Charrington
11 Building AI Products with Hilary Mason - #11
Building AI Products with Hilary Mason - #11
The TWIML AI Podcast with Sam Charrington
12 Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
The TWIML AI Podcast with Sam Charrington
13 Understanding Deep Neural Networks with Dr. James McCaffery - #13
Understanding Deep Neural Networks with Dr. James McCaffery - #13
The TWIML AI Podcast with Sam Charrington
14 Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
The TWIML AI Podcast with Sam Charrington
15 Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
The TWIML AI Podcast with Sam Charrington
16 Machine Learning in Cybersecurity with Evan Wright - #16
Machine Learning in Cybersecurity with Evan Wright - #16
The TWIML AI Podcast with Sam Charrington
17 Interactive Machine Learning Systems with Alekh Agarwal - #17
Interactive Machine Learning Systems with Alekh Agarwal - #17
The TWIML AI Podcast with Sam Charrington
18 Location-Based Intelligence for Smarter Marketing with Klustera - #18
Location-Based Intelligence for Smarter Marketing with Klustera - #18
The TWIML AI Podcast with Sam Charrington
19 AI-Powered Customer Support with HelloVera - #18
AI-Powered Customer Support with HelloVera - #18
The TWIML AI Podcast with Sam Charrington
20 Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
The TWIML AI Podcast with Sam Charrington
21 Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
The TWIML AI Podcast with Sam Charrington
22 Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
The TWIML AI Podcast with Sam Charrington
23 From Particle Physics to Audio AI with Scott Stephenson - #19
From Particle Physics to Audio AI with Scott Stephenson - #19
The TWIML AI Podcast with Sam Charrington
24 Selling AI to the Enterprise with Kathryn Hume - #20
Selling AI to the Enterprise with Kathryn Hume - #20
The TWIML AI Podcast with Sam Charrington
25 Engineering the Future of AI with Ruchir Puri - #21
Engineering the Future of AI with Ruchir Puri - #21
The TWIML AI Podcast with Sam Charrington
26 Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
The TWIML AI Podcast with Sam Charrington
27 Introducing Psycholinguistics into AI with Dominique Simmons- #23
Introducing Psycholinguistics into AI with Dominique Simmons- #23
The TWIML AI Podcast with Sam Charrington
28 Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
The TWIML AI Podcast with Sam Charrington
29 Offensive vs Defensive Data Science with Deep Varma - #25
Offensive vs Defensive Data Science with Deep Varma - #25
The TWIML AI Podcast with Sam Charrington
30 Global AI Trends with Ben Lorica - #26
Global AI Trends with Ben Lorica - #26
The TWIML AI Podcast with Sam Charrington
31 Intelligent Autonomous Robots with Ilia Baranov - #27
Intelligent Autonomous Robots with Ilia Baranov - #27
The TWIML AI Podcast with Sam Charrington
32 Reinforcement Learning Deep Dive with Pieter Abbeel  - #28
Reinforcement Learning Deep Dive with Pieter Abbeel - #28
The TWIML AI Podcast with Sam Charrington
33 Robotic Perception and Control with Chelsea Finn  - #29
Robotic Perception and Control with Chelsea Finn - #29
The TWIML AI Podcast with Sam Charrington
34 Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
The TWIML AI Podcast with Sam Charrington
35 The Power of Probabilistic Programming with Ben Vigoda - #33
The Power of Probabilistic Programming with Ben Vigoda - #33
The TWIML AI Podcast with Sam Charrington
36 Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
The TWIML AI Podcast with Sam Charrington
37 Video Object Detection at Scale with Reza Zadeh - #34
Video Object Detection at Scale with Reza Zadeh - #34
The TWIML AI Podcast with Sam Charrington
38 Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
The TWIML AI Podcast with Sam Charrington
39 Expressive AI-Generated Music With Google's Performance RNN with Doug Eck  - #32
Expressive AI-Generated Music With Google's Performance RNN with Doug Eck - #32
The TWIML AI Podcast with Sam Charrington
40 Smart Buildings & IoT with Yodit Stanton - #36
Smart Buildings & IoT with Yodit Stanton - #36
The TWIML AI Podcast with Sam Charrington
41 Deep Robotic Learning with Sergey Levine - #37
Deep Robotic Learning with Sergey Levine - #37
The TWIML AI Podcast with Sam Charrington
42 Deep Learning for Warehouse Operations with Calvin Seward - #38
Deep Learning for Warehouse Operations with Calvin Seward - #38
The TWIML AI Podcast with Sam Charrington
43 Cognitive Biases in Data Science with Drew Conway - #39
Cognitive Biases in Data Science with Drew Conway - #39
The TWIML AI Podcast with Sam Charrington
44 Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
The TWIML AI Podcast with Sam Charrington
45 Web Scale Engineering for Machine Learning with Sharath Rao - #40
Web Scale Engineering for Machine Learning with Sharath Rao - #40
The TWIML AI Podcast with Sam Charrington
46 Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
The TWIML AI Podcast with Sam Charrington
47 Machine Teaching for Better Machine Learning with Mark Hammond - #43
Machine Teaching for Better Machine Learning with Mark Hammond - #43
The TWIML AI Podcast with Sam Charrington
48 LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber  - #44
LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber - #44
The TWIML AI Podcast with Sam Charrington
49 Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
50 Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
The TWIML AI Podcast with Sam Charrington
51 Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
The TWIML AI Podcast with Sam Charrington
52 Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online  Meetup
Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
53 Word2Vec & Friends with Bruno Gonçalves -#48
Word2Vec & Friends with Bruno Gonçalves -#48
The TWIML AI Podcast with Sam Charrington
54 Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan  - #49
Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan - #49
The TWIML AI Podcast with Sam Charrington
55 Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
The TWIML AI Podcast with Sam Charrington
56 Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
The TWIML AI Podcast with Sam Charrington
57 AI-Powered Conversational Interfaces with Paul Tepper - #52
AI-Powered Conversational Interfaces with Paul Tepper - #52
The TWIML AI Podcast with Sam Charrington
58 Topological Data Analysis with Gunnar Carlsson - #53
Topological Data Analysis with Gunnar Carlsson - #53
The TWIML AI Podcast with Sam Charrington
59 ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
The TWIML AI Podcast with Sam Charrington
60 Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
The TWIML AI Podcast with Sam Charrington

This episode explores the latest techniques for running generative AI on-device, including diffusion models and novel distillation frameworks. Hung Bui discusses his team's work on SwiftBrush and SwiftEdit, which enable high-quality text-to-image generation and editing in a single inference step.

Key Takeaways
  1. Implement a distillation framework to guide the training of an efficient student model
  2. Use a secondary 'coach' network to align the student's denoising function with the teacher's
  3. Apply inference-time scaling to optimize model performance under a fixed compute budget
  4. Deploy high-efficiency diffusion models on-device for image generation and editing
💡 Novel distillation frameworks can enable high-efficiency diffusion models to run on-device, allowing for personalized image generation and editing capabilities.

Related Reads

📰
50+ Sequential Images, One Prompt in Codex
Learn to generate sequential images with Codex using a single prompt and understand the limitations of this approach
Medium · ChatGPT
📰
How can I batch-generate 3D assets from prompts or images using an API, and which 3D generation APIs support batch generation?
Learn to batch-generate 3D assets from prompts or images using APIs for efficient pipeline creation
Reddit r/artificial
📰
How AI Head Swap Works: The Technology Behind Realistic AI Image Replacement
Learn how AI head swap technology works and its applications in image editing
Dev.to AI
📰
How I Built an AI Pet Portrait Generator That Turns Photos Into Art
Learn how to build an AI pet portrait generator that turns photos into art using deep learning techniques and Python libraries
Dev.to · William Li

Chapters (7)

Introduction
4:44 VinAI Research
7:03 Building an AI team
9:02 First AI residency program
9:47 Focus on model efficiency
11:49 Training a Vietnamese LLM
16:29 Model optimizations for Viet
Up next
Topview AI Features EXPLAINED: Pros & Cons HONEST Review
MaxonShire
Watch →