InstructPix2Pix (w/ OpenAI's Tim Brooks)
Key Takeaways
The video discusses InstructPix2Pix, a generative model that learns to follow image editing instructions, with Tim Brooks from OpenAI, covering its applications, technical details, and research implications. Tools and techniques such as OpenAI, InstructPix2Pix, Diffusion-based image editing, and fine-tuning are demonstrated throughout the video.
Full Transcript
okay let's let's start um we have Tim here today uh he's joining us from openi he's one of the co-authors of the of the Del three paper and also previously one of his very exciting Works which he's going to be talking about today is going to be is in instruct P to PS uh before that well he's done a ton of things so I don't know where to start but like you did work at Nvidia um in particular with teroc carus which is the guy behind styan if I'm not wrong right yeah that's right he's he's an amazing researcher as well and like one of the things that caught my eye in particular is that te has a very non-standard background in the sense of being award-winning photographer for National Geographic and those types of things really caught my eye because it's not just like hard tech and AI but you also had this let's call it soft side or like artistic side to you which is super cool uh there is not that many people who who kind of have both of those two I I know that Paul Graham for example is a very prominent example of somebody who went to I think arts school as well and then computer science so I find that combination fascinating and plus you are literally working on the lead three which is again synthetic photography or like exploring the latent space and just like kind of um decoding that into the image space so without further Ado I think team you can K it off I'm GNA start stop rambling on my side too much coding too much debugging awesome uh thank thanks so much for having me let me present hey guys I want to give a huge shout out to hyperstack folks who have generously sponsor my compute over the past month or so uh basically I got 16 h100s which is two uh e GPU notes and the performance has been amazing basically two exra speed up compared to my a100 notes so I want to quickly show you how you can get started yourself it takes basically three steps only you go to the environments here you create new environment you give it a name you pick between Canada and Norway that's the First Step the Second Step you go to uh SSH Keys here you create a new pair you pick environment you just created you give it a name you paste your public SSH key that's it and finally go to Virtual machines deploy new machine give it a name again select environment let's select Canada here because they have more compute there then you select the hardware you want like for example h100s you select the images you select the SSH key you just created and hit deploy that's it literally a couple minutes to get started it was really easy for me to to just uh basically create these two nodes I just mentioned and get started running llm trainings so the documentation is super cool I I could actually solve all of my problems just by looking at their Ducks they additionally have a slack Channel where they were super helpful so I can't recommend them enough uh honestly the the thing they kind of focus on is basically because a lot of these GPU providers focus on on big Enterprises so often times you can't get on demand h100s whereas next gen focuses particularly on that you can get top of the line Hardware on demand even if you're if you're individual or like a smaller team or whatnot they also focus on bigger companies but that's kind of there their edge here so without further Ado guys let's go back to the talk I do suggest you check them out and let's continue can you all see my slides now yep we can see it awesome cool so yeah I'm I'm Tim I'm a research scientist at open aai um I'm going to talk today primarily about uh of mine instruct picks to picks um and the main angle that I want to talk about it from is about the methodology in that work being about the data and actually in a lot of Works um that I've been a part of and I think in a lot of the most impactful image generation work uh data being the main method um and I have some contact info here feel free to like reach out I'd be happy to chat about any of this um afterwards as well so a bit on my background um as you were just mentioning I actually come first from a photography background and from when I was a kid to kind of like my first jobs that I was doing was all in photography and um I really loved creating images so here's some of my artwork mainly nature photography um and it was always really technical for me um about how to push the limits of cameras and really understand understanding cameras well and trying to leverage the technology to capture moments that people don't normally get to see uh and that led me to my first job of trying to make that more accessible for people who don't have the experience or the equipment uh to be taking these photographs if we can make it so that on normal smartphone cameras people can take really amazing photos so that is uh I joined Google AI and I worked on a research team that was doing AI for computational photography specifically on the pixel phone camera and I'm just going to quickly go through a little bit of that uh before talking about instruct pix to Pi as the main work and so we we did a few things one technology I was part of was making uh night mode on the pixel 3 phone and at the time that was really exciting because it was one of the first times you could really take pictures in very dark settings uh so that was really exciting to to announce and then moving forward we even push that to being able to take uh photographs of like stars it's extremely dark um and this other technology I worked on ended up launching for for Action photographs and for creating motion blur and these are all kind of the theme Here is like trying to allow people to take photographs that previously would have required professional equipment and maybe professional expertise um and the actual research that I was doing to make this happen was really all about creating training data and one of them was on D noising and there are a ton of papers on D noising and at the time a lot of people were working on improving d noising by improving architectures some new layer that you add or some slightly new architecture to your convet or whatever it was um but in this first work I did at Google it was all about just making the training data really really accurate and high quality for D noising and that helped so much more than anything else we did in making AI based noising actually helpful in phone cameras uh and this next work was on synthesizing motion blur and this again was all about creating the training data and what I did here was I actually used one neural network that was really expensive that you had to run a ton of times to create really high quality synthetic training data for another model that was a lot faster and that did this problem directly so both of these Works were about creating training data and they were about trying to unlock these Technologies uh that were previously kind of inaccessible unless you were a professional and around this time I got really interested in okay it's cool if we can make photography more accessible but there are so many images that people might want to make that even if you had a camera and the best experience as a photographer you just can't create they're only in your head and that's when I got really interested in generative modeling and I went to UC Berkeley uh and did a PhD in generative models and I was advised by alosa OS there and one of the main works that I did is instruct pick tops learning to follow image editing instructions and that's what I will be talking about today awesome this is this is super cool before you go into the technical topics can you go back to the first slide that kind of took my attention caught my attention with the turtles and the fox like did you actually so so you were basically diving when you took the photo of the turtle can you give a bit more background there like I think that's kind of this part is super interesting just just to put it like yeah everything else we gonna talk about a bit later totally I mean that the most amazing thing is like actually being there especially like in the upper right being by that coyote was just one of the coolest experiences I've had hands down I think um but yeah the turtle was in in Hawaii um I was like snorkeling with you know a camera and trying to get this picture of a turtle um and all of these were um a lot of it is like the amount of time that goes into one image is really high and it would be like you know you spend a day tracking a coyote and waiting for the right moment and setting everything up um or just hours waiting for a bir to land to get a photo of it um now replace that with bugs now replace that with bugs and you're literally an open eye that's it that's it yeah yeah you spend a whole day soling a bug yeah exactly I mean I mean it does require similar type of like perseverance being being patient and and tracking stuff so like from that very high level like a standpoint there is some comparison to be made but like isn't kote like dangerous or something you were literally going around Wasteland there and taking photo of of wild animals like what what's the thing there yeah I I mean I do think it can be a bit dangerous you need to like read up a bit on wildlife and how to be kind of cautious and an important thing is to make sure that the animal that you're not like approaching it headon and that you're really gradual and slow and you're down at its eye level um or maybe it's even like coming toward you so you give it the option if it wants that it doesn't need to continue coming toward you super difficult okay if other people have question feel free to do that if not the the next slid you show with pixel I'm curious like P 10 were you working much much more on the on the research side developing the architecture of like the Comets or whatnot or were you also on the on the production side like just shipping this to actual pixel phones on the on the edge because both are kind of very exciting yeah one could be working on this project I'd say it was like 80% on research and 20% on deployment um which is kind of a balanced that I actually really like um but yeah more focused on the research and like how can we make cameras and two years from now or five years from now be extremely good uh but it's fun to have this like a little bit of like let me get something out there right now and see people use it yeah I guess the exciting part here for me is in particular if it's actually running on the pixel phone as as On The Edge versus just you sending request to the API that's kind of exciting because you probably had to do a ton of Interest to work on distillation and quantization or whatnot to to Shi yeah making them really fast is hard and I think that's actually another very strong motivation for formulating the problem in terms of getting the perfect training data because once you have the perfect training data then it's easier to make sacrifices on making your model kind of smaller or more efficient um that can work on the device because you don't have the overhead of learning from imperfect data super cool awesome I think you can you can now continue as far as I'm concerned cool all right so instruct picks to picks learning to follow image editing instructions here was kind of the goal is to be able to interact with an image editing model in this conversational instructional way where you give an image and an instruction for how to change it could you make this image look like it's in time and just responds from that instruction so this is the motivation here are a bunch of results of what the final model was able to do so you can like swap objects add something in the background you can pose the instructions in many different ways like you could pose it as a question what if it would look like it was snowing uh you can change specific attributes like making the jacket out of leather uh so this is what we were building and I think there are a number of ways we could have thought about this problem formulation uh but one that really made sense to me is that if we had the perfect data for this task I actually think it's a pretty easy problem to solve if you have a ton of pairs of input output images and the instruction that goes along with the change between the input and output you can just train a generative model to do it it's like a supervised learning problem and this reduces the problem entirely to that of how how do we get this training data and That Was Then the focus of the work uh also really important to kind of contextualize it into the research that was going on at the time so there were a number of papers on diffusion-based image editing uh they were looking really promising uh a lot of them involved this process of fine-tuning a model on each example or like dream Booth which is really cool work you f tun on a collection of some images and then you have this model that's specific to that concept that's learned and you can create new images based on that concept um or like a magic happen around the same time it also had this adaptation and fine-tuning and optimization for the particular image but one thing that I think is really important for having it be conversational in the way that I was imagining is to have it be really fast so you don't require this per Image level optimization um and also just trying to formulate it in some way that it happens directly is really nice to just simplify the problem rather than requiring this like per example optimization or extra inputs just is there a way we can get a model to do this directly as the inputs and outputs and another related work was prompt to prompt and this is a bit different but it takes in two different captions and it generates two images so both of these images are generated of the cat um but it's able to make sure that the two gener generated images look similar and it does that by copying the intermediate attention maps from one image to the other while it's generating them and that forces the spatial layout or structure of things uh to look similar between the two different examples uh but it still has the difference in the prompt so you can see whether it's a bicycle or a car and so this idea is really important because this is what we use uh in order to generate training data and so here's just showing that interactive use case um that we were Imaging again and another important but subtle thing is the actual phrasing of the language being able to get it to work with instructions rather than um requiring someone to describe the entire input image for example and the entire output image um or provide extra conditioning images these are the only inputs and outputs we wanted and I guess now maybe this looks more obvious I think uh since chat GPT has come out that this is a really great way to interact with AI models um at the time which was like only a month before chat gbt this type of conversing back and forth uh was a bit less common um but I think it's something we should strive for for like all different modalities definitely for image generation and image editing as well and so the approach that we use is to just to train TR a large diffusion model directly to edit images and train it on a large supervised data set of paired images and instructions maybe a quick question here before you go U because you mentioned on the previous slide that the chat format being kind of common way currently because of jgpt but like going going forward do do you see do do do you see any interesting work on just how how we interact with these systems or you think like that chat format is is here to State yeah it's a it's a good question I think we will always be hitting the limits of how we interact with them and like eventually we'll probably have mind reading and then maybe we'll really be able to get models to do exactly what we want but until we have that I think there are always going to be these boundaries between what we're imagining and the interface for interaction so I bet there will be a lot of cool Works in combining chat with like maybe other types of interaction action for trying to edit images um and combining them so I I I don't think chat is like the only you know this allows a lot of cool image editing and interaction but there are things you can't do well with this too um so I I yeah I think that we need um a combination of things we need mind reading got it yeah we need mind reading we actually had to Nish some weeks ago where he was presenting one of his Works where basically they took the fmr fmri scans and then decoded that into image space so like like imagine fmri being something that's like varable device at one point in the future which doesn't seem impossible like that that type of work would be super cool like you just imagine something and then like literally reads your mind and you're all of a sudden in the image space like that that would be super totally this was actually one of the things that I was like considering doing my PhD in this area and I totally believe that it it will happen I mean it will be hard to make the machines small enough and High Fidelity enough but at some point it definitely will happen and we will just be able to like visualize our dreams which will be so amazing and scary at the same time right yeah yeah with every technology comes the comes the misuse but totally cool yeah all right so so the main question now is where does this supervised data set come from from and the main idea is that we're going to use these large pre-trained models that we have gp4 and stable diffusion are the main ones that we use but we're going to use this immense knowledge or sorry GB3 these immense knowledge that these large models have in order to generate training data with them and so here this is again pre- chat gbt days this is the playground for using gbd3 um I wrote out some examples of an input image caption and an instruction for editing it and then a corresponding output if that edit had been applied so for example an image of a person holding a cup of coffee and the edit is turn the cup of coffee into a bowl of soup and then the output would be an image of a person holding a bowl of soup so these are examples of image editing but entirely in text space you don't actually have any images here and the cool thing is that by learning from the context of a few examples GPD 3 can then generate the correct output so here landscape photograph of a lake with mirror likee Reflections green summer trees and if you say the edit change the season to Autumn and you just ask gpt3 what text comes next it will generate um landscape photograph of a lake with mirror like reflection autumn leaves so it understands how to create that to edit and we can even take it one step forward by just giving it an input description a young man with brown hair wearing a green backpack and GB3 can generate for us both a plausible edit give him a blue backpack and the actual output a young man wearing brown hair wearing a blue backpack so so this is really the main driver behind the intelligence of what an image edit means is that we're going to use GPD 3's knowledge in text space to create synthetic training data but only in text space and instead of just uh giving it three examples in context uh we actually wrote 700 examples ourselves which was a pretty grueling process um but then you can find tune gbd3 and if you find tune on like 700 examples we found it works just extremely well at this specific test um so we took some captions from The Lion data set we manually wrote the edit instructions and the edited caption ourselves for 700 of them and then we use the finetune model to automatically generate the edit instructions uh of 450,000 more and that's what's in green here these are generated with GPT quick question now because you mentioned 700 edits and you're mentioning gpt3 um um given the context length of gpt3 like I'm confused on that on that part there yeah yeah what do you do with 700 prompts at the same time we actually ftuned I fine tuned gbd3 there was like an API for fine-tuning where you can upload paired text examples and it does like I think like it did two epochs of training on it and just a little bit of fine-tuning okay that makes sense I I thought it was few shot but that's it like too much too much even for yeah that's right I used F shot in the first example and F shot was like the first thing that I tried to kind of realize that this would actually work but then um fine tuning it works works better yeah super cool I think prun has one more question prun go ahead hey hello can you hear me yeah yeah cool uh okay so how do you validate the data that you find tune because 400k uh Rose is like how do you like how do you know it's true to the human uh data set yeah that's a great question so what I what I did was and I tried for example like I first did it on not 700 I first tried just writing 100 and I tried fine tuning it and I just looked at like I don't know how many random samplings of a few hundred and like scan through to see how accurate it was and it wasn't working super well which is why I then wrote 200 and then you know eventually 700 and I just kind of manually inspected like order of a hundred of these um and at some point decided that they were they were looking pretty good they definitely are not perfect like if you look through sometimes the fine tun gpt3 just makes an error but also the input like Lion is a kind of noisy data set um so sometimes those captions would be pretty bad too uh and it kind of got to a point where gbd3 wasn't limiting things any more than the quality of Lon and once we hit that point um we were in a in a good spot to move forward super Co okay okay cool thanks y on this slide one more question about the edit instructions because we have 450,000 there like how did you generate the edit instructions then the second part is clear once you have this fine tune model yes so it actually like when we go here in this example it generat gpt3 is generating both the edit and the output I just gave the example of it a young man with brown hair wearing green backpack and so given examples of it GPT can understand what a plausible edit would be and it can create one and that's what we do here uh when training on it the input lion caption is is the input to gpt3 and the output example that we have it generate is the concatenation of the edit instruction and the edited caption and then when we run the model we give it just the input lion caption and it generates again the concatenation with a like a special token in between to to Signal the the break it generates both the edit instruction and the edited caption for us sense thanks and so now that we have all these pairs of before and after captions we can use a text to image model uh we use stable diffusion to turn that into training data to generate images from them so we use that prompto prompt method that I mentioned before prompto prompt helps ensure that the two images look similar to each other um and this is how we're going to turn the text Data into image data so just as a quick recap of all of this the overall process for creating a training example is that we have an input caption for example a photograph of a girl writing a horse that goes through a fine tune version of gbd3 and the fine tune GB3 produces both the instruction have her write a dragon and the edited caption photograph of a girl writing a dragon we then take the input caption photograph of a girl riding a horse and the edited caption photograph of a girl riding a dragon we pass both of them through stable diffusion and use prompt to prompt to turn them into a pair of before and after images and that constitutes one example in our training data set and we do this 450,000 times now we have a really large data set uh exactly for the task that we're trying to do of an input and output image and the instruction that corresponds with that and one thing that's important to mention that this slide doesn't really show that the data is actually really noisy it definitely doesn't this process doesn't work 100% of the time sometimes you know as we were just talking about sometimes gpt3 fails or the lion caption might be bad sometimes stable diffusion has a failure or prompto prompt has a failure um but we generate enough of this that even though the data is noisy there is signal to learn from in there and we also uh use a Critic we used clip to help filter out the bad examples so instead in of just generating one for each of these for each um caption pair and edit instruction we generated a bunch of different candidate image Pairs and then we use clip to figure out which of these uh examples is the highest quality and we filtered the data and I think this like the main point that I want to make about this work and just in general is like that is really the method everything that I covered so far is really the method there's also a little bit on top about like training a diffusion model and you know how the input goes in and a bit about how the conditioning and the classifier free guidance works and the hyperparameters for training it and such but really the process of generating the data is the main method and I think just as a as a general rule um when we're trying to solve really hard tasks instead of spending a lot of energy making like a complicated model that has these specific modules for how to do it our time is better spent just putting all our energy into making really great training data um and then training the simplest model that we can on that data and a couple extra reasons why why I really like uh this approach um well one is that making great training data I actually think is a lot easier than working with bad training data working with bad training data is really hard it's like really hard to force a model to overcome some limitations in it data so actually think this is the easier approach um and it's also useful for a long time and for many models so with an instruct picks to picks the way that I used the data in the paper was to uh have a fine-tune version of stable diffusion but we also released the data and a bunch of other people have then used the data for like new models that they're developing as part part of larger projects or in specific ways and control net with improvements and so having data as this asset that we share I think uh can actually Outlast uh the model itself and be really useful hopefully for for a long time uh and now this is that like extra part okay we we do actually train a model on this uh it's just a supervised learning problem now but what we do is fine tune stable diffusion on this generated training data and we add zero initialized image conditioning channels because now we have an extra input uh stable diffusion normally just goes from noise to image but we have this conditioning image as well so we add um some extra channels in order to condition on that and something that's really amazing is that after all this the model generalizes to take in real photographs in real paintings and this was something that we weren't even sure if this would happen when we were doing the whole project leading up to here I mean certainly I really hoped that it would um but all the training data is entirely synthetic even the text the text is generated by GPT not by written by people and all the images both the input and output images are all generated U by stable diffusion uh but it amazingly is able to generalize like a real painting or a real photograph I mean this definitely seems to be a common Trend I I remember reading maybe three four years ago your paper on um solving Rubik's Cube using the robotic hand the textas hand and if I if I recall correctly even then the finding was if you have synthetic data which is not even maybe of that high quality but you randomize it a lot like for example you modify gravity or whatnot and you just generate a [ __ ] ton of data like the model did manage to solve actual Rubik's Cube in the in the real world so like that that definitely seems to be a trend like if you have enough data even if it's synthetic and it's it varies a lot like it boils down that the real world becomes interpolation of what the model has seen almost interpolation ideally of what the mod has seen during training right yeah that's I see it yeah I think that's right it's a really powerful kind of General thing and it it's important that the data is good enough there and you know there have been attempts to do this where the data wasn't yet good enough and then it didn't generalize so it needs to cross some threshold of quality and scale and diversity of training data um but once you reach it it it's kind of amazing there also there's this really cool work from uh my friend Ilia at Berkeley that just did something similar with humanoids and it has simulated uh robots um and then just zero shot it's able to have real humanoids walk around in the in the real world and so like that I think is really amazing that it it generalizes to that real world use case but I bet we'll see this more and more and that approaches for creating data will just get better and better super cool and I definitely see a trend here as well like J James was also super bullish in doing his talk on on like just data data data generating more data and like yeah and um so in that sense I see I see I some commonality between people working on3 and and it makes complete sense ultimately like model if it's if it's like high capacity enough and maybe maybe the efficiency bit is maybe important because lat diffusion models obviously made it much more efficient to train yeah but like other than that you have a capable model that can just like learn the underlying function you give it a ton of data that's it yeah I think that's right and and this isn't to say that there aren't aspects of modeling that are also important and like you know getting good scaling properties and models and and doing work there I think is really valuable as well um but I think data has historically been undervalued and really when you just look at the relative importance of things it is so so so valuable 100% I think andang was kind of maybe a Pioneer in in that regard correct me if I'm wrong but like he had this um data challenge or something where where like the idea was like you keep the model fixed and the whole idea make the better data set and challenge if I if I recall correctly some years ago maybe yeah that that is a it's a really cool way actually to try formulating progress and I think part of the challenge is about just how Publications work in Academia that there's a lot of incentive to have new methods um but posing a challenge specifically that way is great because it also creates this incentive of like an actual um like challenge with a leaderboard about data yeah yeah yeah 100% I think he called it like data Centric something approach or yeah data Centric AI yeah yeah yeah yeah yeah and and and and briefly comment on the on the latent diffusion arguably the the the biggest advantage of that method was that everything was happening in the Laten space so that's efficiency so bit bitter lessons all over again I like yeah I don't know whether one could argue that that was the main the main thing yeah um yeah if you and you even you kind of can to some extent look at it from the standpoint of How It's modifying data too because what latent diffusion is doing is it's it's changing your data instead of being in pixel data it's now latent data and latent data is smaller and more efficient so it's a an optimization in some sense of your data to make it easier to model yep awesome thanks all right so here here are just some fun examples examples girl with the Pearl Earring you can edit a bunch of different ways uh the instructions are below and here are some real photographs um adding boats or adding a city skyline making it different times of day so it really has like a wide diversity of types of edits which I think is also important there uh were a lot of Prior works that kind of were specialized like just doing style transfer just doing one thing so it's a nice property that you can zero shot do all these different tasks and one uh technical detail that was really important to get this to work well is the use of classifier free guidance um for the two different conditioning signals so classifier free guidance uh for anyone who's not familiar is a common method with diffusion models where you make the conditioning more strong you do this by extrapolating the score which is the output after each G noising step uh toward the direction of having more conditioning and people do this for text to image models we also did an instruct pick to pick to make the instruction stronger you can see the effect here when you use just a value of three versus 15 for turn him into a cyborg with a higher value it um appli the edit more and more strongly but something that we found was important because we actually have two inputs now we both have text and we have the input image um is to extend this uh to use different scales for these two different inputs one for the image and one for the text and it was really important that we separate the two scales uh because they weren't calibrated if you use the same scale for both of them um you did not get great results and you can see on the on the right most or on the left on the column that the scale that we use for images is much smaller but it does help you get kind of the best results when you use some combination of the two and you use stronger instruction conditioning as well as stronger input image conditioning um so this is kind of a technical detail but it was a really important one for getting good results and to get the very best results uh you can even find to you can tune these manually for a particular example to get the exact kind of balance of how well it adheres to the input image versus uh how well it follows the prompt and so the experiments that we did uh to make sure that this was really uh impactful was uh just mainly about data too like does the data scale and data quality help and we have some some metrics here where kind of the outer outermost curve is is the most preferred and so data helps a lot um I'll go through some of these a bit quickly but we compare a prior prior methods um and it also was you know quite the step above the existing work at the time you can generate a bunch of different um varieties of edits um just by sampling different input noise uh one thing that's nice about it is it generalizes to high resolutions well and this works particularly well because it has an input image which kind of gives the structural layout um you can actually run it really high resolution and at different resolutions um it's fast which enables this iterative editing or like doing it in the chat uh interface one challenge with it um which is just generally really important and challenging with image generation models is there's lots of bias um and so this example I think demonstrates it well when we ask uh to make them look like flight attendants it makes them more feminine and make them look like doctors makes them more masculine um and I think that what really be desired is to keep the identities of the people the same and and just change the appearance to make those same people um look like a flight attend or doctors rather than actually changing their identity in this way so that's an important thing to try improving is kind of disentangling these these correlated and biased factors and it also has some failure cases so particularly it can't do structural changes well like zooming into the image um or this one moving to Mars it like doesn't isolate the object so well of the Eiffel Tower uh color the tie blue it doesn't isolate and I think a lot of these limitations are actually limitations I mean they are they're limitations of the data and I think the way to address this and the way to address the bias from before is by improving the data these respects improving the methods that are used to generate them whether it be stable diffusion um or prompt to prompt or the GPT model that was used to create these prompts and yeah people have created a lot of really cool uh applications and things with instruct Basics which is which is nice and some of my main takeaways from it are that uh image generation models can be made more ful by making them instructional and conversational um rather than just having it uh Tak in like prompt C captions that require these maybe lengthy and and very like prompt engineered inputs uh by just being able to converse with image generation models um and this other one is like that we can use the existing large models that we have whether it be gbt or stable diffusion or other models and other modalities that we can use them to create these really large uh pre-training data sets for other more complicated tasks even tasks that maybe one model individually can't do which was the case here like neither GPT nor stable diffusion could create this data on its own or do the task on its own but by combining them we were able to create the training data for it um and that's just again this main point like data is the main thing it's the method um and that that is all I have for today um so here are some of the main results again awesome definitely much more images than the James you won this [Laughter] challenge um even has a question even go ahead uh okay sorry I joined a bit late uh but I wanted to ask uh why do you think these models are so bad in text like uh I mean they can do large text by now like if you have small text on image but why do you think they can't Define gr text yeah totally um it's a good question I think that I mean text is very special to humans also human faces are but text in particular like when we see characters they have this really specific symbolic meaning to us and in images when you're like the model just doesn't yet know that and it needs to transform from this conditioning which is like very different the conditioning is just like this embedding Vector that maybe corresponds with like a token that's a discrete value of these characters to something like totally different which is this rendering of a character it's not like in a language model it's like the representation that's input and the representation that's output is the same they're both these discrete tokens that correspond with characters but here just the inputs and outputs are so different that I think it's challenging for the model uh to do that well and also we're just like really really we notice the issues whereas here like if you look at this Eiffel Tower right like it totally has issues on the the textures or the details of things but like that's not really a big deal I mean it'd be nice if it weren't there but we have these imperfections in other things um but they're not a huge deal if you like repeat a texture in like foliage or something twice when you shouldn't um so I don't think we notice those as much but we really notice if you repeat the same character twice so I think part of it is that text is very important to us and the model doesn't know that it's any more important than you know other imperfections that might create but do don't you think that it's like indication of larger problem like uh do you think that it's because of embeddings being too small and they that you have a single embedding as opposed to multiple embeddings in text models yeah that's a good question um I mean I think it's totally possible I don't really know I think that would be like a cool research idea to try experimenting on um but it's yeah it's totally possible that like the representation that we're embedding uh for these texts is limited and that by having larger or multiple or different embeddings that maybe would solve these text rendering issues um yeah it would be cool to try okay I think one more question from Sun cult yeah hi Tim hi it's lucky to meet you so my question is that uh uh during this experiment or this journey did you uh did you uh try to visualize this uh uh this uh what you say learning to edit and how did it look did I try to visualize the learning to edit means uh at every instructions it was learning it was learning to edit like how the cat sits on the chair and then suddenly you change the background and the same alignment gets fixed on something else so basically model has somewhere learned the sting I guess the art of sitting and uh yeah so did you try to visualize I I guess the only visualization that I did was like by editing images um but if you mean something about like visualizing internal representations um or like something in that vein I I didn't try that at all um and yeah that too could be I know there are some like approaches for visualizing maybe what uh certain promps like most correspond with or activations or something if that's what you mean um but kind of the the main visualization I used was just running the model and seeing how the edited look yeah that's also great yeah it's actually understanding like the replace the fruits with the cake the the fingers and the cake base is totally uh means uh settle like that's that's it's beautiful yeah yeah thank you um how I understood his question how I understood it was maybe that he was asking about gpt3 as well like the types of edits that the model makes and somehow clustering those into categories like that that's how I understood his question maybe more in the text space before you even go to the image space yeah thanks Alexa it got it um no yeah I didn't do and and what are there particular ideas for how to visualize it that that you have for like so we we have this large data set of like 450,000 um all I did was like read through examples of it but do you have ideas for for what types of visualizations of that you think would be interesting are you asking to me or him I guess either I mean yeah it would be cool to like clustering is a idea I totally bet there are like modes of like replacing objects is probably like a common mode of types of edits yeah if you like uh what I feel if you like to Cluster then you will know that where it has learned to sit uh in a more better way me uh the editing thing is like you are trying to settle in the environment that's what you say edit otherwise you say it's a uh it's it's a fault in a like a class something like that me it does not fit to the uh the scene yeah like for example we know we know that one cluster that's probably underrepresented or doesn't exist is the cluster that basically modifies the structure of the image significantly right because you mentioned that's one of the filler modes in the image space so like likely if if you give it the image of Eiffel Tower it's going to modify the color the sky color the texture stuff like that but it won't say like I don't know like morph this this Eiffel Tower into a like metal cup or something like that I don't know to totally totally and I I bet that if you did this visualization clustering you'd find that where you're near lots of data points the model does well at editing and any of these like gaps like that example exactly where the model does that yeah yeah know that's a that's a cool idea super cool can you go back to the application slide you kind of schemed over it but I think it's very important and it might be fun for a lot of people like what are some of the like most interesting applications you've seen for example the the way I found you by the way was like actually going to the top hugging face spaces I was doing some some scraping and stuff and like I noticed your your your repo and uh so like it's one of the most popular repos on hugging face like can you maybe walk us through some interesting examples you've seen so far with with instruct PX to pics totally um I mean so one thing that was really fun is like people trying to make these actual uh like texting interface in the same way of that like I just kind of didn't actually hook it up to iOS or anything I just manually made that visualization but that was really cool to see because that is kind of an image that I had in my head for a long time of what a cool application would be um but also in like playground AI has instruct pick to picks in it and I think just it's really cool to see people actually use I mean it's great to like make technology and try making research just to like Advance the field or understand something new um but I think this was like the coolest part of this project for me was made it and put it out there and the fact that people like made applications around it and started using it it was just like so cool to see uh that people actually want this technology super cool how much are you leaning towards productionize research you mentioned kind of 2080 is aweet spot but like how do you feel how do you feel I guess I guess my question is also we had this weird dynamic over since 2022 where a couple of researchers create something that literally opens up a huge Economic Opportunity for so many companies like arguably arguably the hund something million that stability AI raised would not have happened without three guys coding in heidleberg lab and doing the lat diffusion models and and whereas those guys probably captured just like a small bit of that financial gain that that that was C happening the industry how do you how do you feel about that Dynamics as a researcher yourself producing all of this and sharing with the world yeah totally that's a it's a it's a good question and it's challenging because I do love doing research and I think that going entirely on like focused on making products with things and like is it's hard to still do that and do research um I really do like this kind of I find like 80% on research and 20% on product because it allows me to like see research through to the to the extent that like people use it um and like at open AI obviously like Dolly 3 the work that our team and that James did there and like getting that in chat gbt is so cool because the research itself and and the paper and the understandings are really neat but the fact that people like actually use them and do things with them is kind of amazing um so yeah I guess it's it's some amount of like a balance but I love research so much that to me like I still want to be putting the majority of my time into that um and I just also want some amount of like trying to get it in people's hands makes sense thanks uh sunclub has the last question and then I think we can wrap it up unless anybody else has a question how do you convince investors like uh when you when you're going for the research because they they go for more like business oriented goals and research oriented is more like uh like you want to do it yeah I mean I think that's tough and like I I don't have a startup so I don't have experience with doing this um but I I definitely think that's a challenge I mean uh there are are some successful cases but often it is the coupling of the two and then sometimes it's the story or idea that in enough time there will be product value even like you know open AI kind of did was a very special case but for a long time they didn't have products and it was really this promise that the research will be so amazing and will and it ended up happening that it ended up creating products that were valuable um so there needs to be some path to it that people can see in order at least through like traditional Venture Capital uh to get funding for that research super cool uh harad uh you can you can ask the next next question hello Tim yes so my question would be during your research did you find there is any need for any successor of diffusion based model during your translation of computational photography and instruct pix to pix or do you feel like did we hit just hit the wall there is no successor like that I I don't think we've hit a wall um and there's a lot to come right I think there's a lot of like it's always what happens is that sometimes particular approaches start to Plateau but I don't think that's like you know we could get these models to do so much more complicated of tasks and the images could look a lot better and there are other modalities and higher resolution and there are like lots of things and what about making it just like insanely fast and insanely cheap um and I think that there's a lot of room um to still kind of push uh the generative modeling approach that's used that said I do think that often the lowest hanging fruit and the most F thing to do is on data which is kind of why I like to put my emphasis there um but I bet just how like you know everyone was using Gams and then you like people got diffusion models to work and they were so good I bet that there will be some really large advances as well in the underlying generative models that are used there are just too many sigmoids going on in parallel that are just stacking on top of each other for for this for this famous fall that's being mentioned completely in the it's crazy that we had so many mentions of the wall like in the period where we we've never had this boom of of AI technology as we had in the last two years and like yeah it's kind of funny yeah yeah I mean we'll have to see it's really hard to predict the future but I'm pretty bullish that that these will get a lot better and that people make some some large research advances that sounds better I mean everyone is working on diff diffusion based models right no one is coming across and and putting their words into new kind of generative model that do the same things but in less time and space yeah I mean certainly a lot fewer people are there are some people doing research on other
Original Description
Become a Patreon: https://www.patreon.com/theaiepiphany
👨👩👧👦 Join our Discord community: https://discord.gg/peBrCpheKE
We had Tim Brooks join me to discuss his PhD work - Instruct Pix2Pix. Tim is also a co-author of the recent Sora work so we might have him soon again. :)
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
https://www.timothybrooks.com/
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
⌚️ Timetable:
00:00 - 01:34 Intro
01:34 - 03:21 Hyperstack GPUs (sponsored)
03:21 - 59:18 Instruct Pix2Pix
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
💰 SPONSOR
The AI Epiphany - https://www.patreon.com/theaiepiphany
One-time donation - https://www.paypal.com/paypalme/theaiepiphany
Huge thank you to these AI Epiphany patreons:
Eli Mahler
Petar Veličković
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
💼 LinkedIn - https://www.linkedin.com/in/aleksagordic/
🐦 Twitter - https://twitter.com/gordic_aleksa
👨👩👧👦 Discord - https://discord.gg/peBrCpheKE
📺 YouTube - https://www.youtube.com/c/TheAIEpiphany/
📚 Medium - https://gordicaleksa.medium.com/
💻 GitHub - https://github.com/gordicaleksa
📢 AI Newsletter - https://aiepiphany.substack.com/
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
#instructpix2pix #timbrooks #openai #sora
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Aleksa Gordić - The AI Epiphany · Aleksa Gordić - The AI Epiphany · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Intro | Neural Style Transfer #1
Aleksa Gordić - The AI Epiphany
Basic Theory | Neural Style Transfer #2
Aleksa Gordić - The AI Epiphany
Optimization method | Neural Style Transfer #3
Aleksa Gordić - The AI Epiphany
Advanced Theory | Neural Style Transfer #4
Aleksa Gordić - The AI Epiphany
Anyone can make deepfakes now!
Aleksa Gordić - The AI Epiphany
What is Computer Vision? | The Art of Creating Seeing Machines
Aleksa Gordić - The AI Epiphany
Feed-forward method | Neural Style Transfer #5
Aleksa Gordić - The AI Epiphany
Alan Turing | Computing Machinery and Intelligence
Aleksa Gordić - The AI Epiphany
Feed-forward method (training) | Neural Style Transfer #6
Aleksa Gordić - The AI Epiphany
What is Google Deep Dream? (Basic Theory) | Deep Dream Series #1
Aleksa Gordić - The AI Epiphany
Semantic Segmentation in PyTorch | Neural Style Transfer #7
Aleksa Gordić - The AI Epiphany
How to get started with Machine Learning
Aleksa Gordić - The AI Epiphany
How to learn PyTorch? (3 easy steps) | 2021
Aleksa Gordić - The AI Epiphany
PyTorch or TensorFlow?
Aleksa Gordić - The AI Epiphany
3 Machine Learning Projects For Beginners (Highly visual) | 2021
Aleksa Gordić - The AI Epiphany
Machine Learning Projects (Intermediate level) | 2021
Aleksa Gordić - The AI Epiphany
Cheapest (0$) Deep Learning Hardware Options | 2021
Aleksa Gordić - The AI Epiphany
How to learn deep learning? (Transformers Example)
Aleksa Gordić - The AI Epiphany
How do transformers work? (Attention is all you need)
Aleksa Gordić - The AI Epiphany
Developing a deep learning project (case study on transformer)
Aleksa Gordić - The AI Epiphany
Vision Transformer (ViT) - An image is worth 16x16 words | Paper Explained
Aleksa Gordić - The AI Epiphany
GPT-3 - Language Models are Few-Shot Learners | Paper Explained
Aleksa Gordić - The AI Epiphany
Google DeepMind's AlphaFold 2 explained! (Protein folding, AlphaFold 1, a glimpse into AlphaFold 2)
Aleksa Gordić - The AI Epiphany
Attention Is All You Need (Transformer) | Paper Explained
Aleksa Gordić - The AI Epiphany
Graph Attention Networks (GAT) | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
Graph Convolutional Networks (GCN) | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
Graph SAGE - Inductive Representation Learning on Large Graphs | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
PinSage - Graph Convolutional Neural Networks for Web-Scale Recommender Systems | Paper Explained
Aleksa Gordić - The AI Epiphany
OpenAI CLIP - Connecting Text and Images | Paper Explained
Aleksa Gordić - The AI Epiphany
Temporal Graph Networks (TGN) | GNN Paper Explained
Aleksa Gordić - The AI Epiphany
Graph Neural Network Project Update! (I'm coding GAT from scratch)
Aleksa Gordić - The AI Epiphany
Graph Attention Network Project Walkthrough
Aleksa Gordić - The AI Epiphany
How to get started with Graph ML? (Blog walkthrough)
Aleksa Gordić - The AI Epiphany
DQN - Playing Atari with Deep Reinforcement Learning | RL Paper Explained
Aleksa Gordić - The AI Epiphany
AlphaGo - Mastering the game of Go with deep neural networks and tree search | RL Paper Explained
Aleksa Gordić - The AI Epiphany
DeepMind's AlphaGo Zero and AlphaZero | RL paper explained
Aleksa Gordić - The AI Epiphany
OpenAI - Solving Rubik's Cube with a Robot Hand | RL paper explained
Aleksa Gordić - The AI Epiphany
MuZero - Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model | RL Paper explained
Aleksa Gordić - The AI Epiphany
EfficientNetV2 - Smaller Models and Faster Training | Paper explained
Aleksa Gordić - The AI Epiphany
Implementing DeepMind's DQN from scratch! | Project Update
Aleksa Gordić - The AI Epiphany
MLP-Mixer: An all-MLP Architecture for Vision | Paper explained
Aleksa Gordić - The AI Epiphany
DeepMind's Android RL Environment - AndroidEnv
Aleksa Gordić - The AI Epiphany
When Vision Transformers Outperform ResNets without Pretraining | Paper Explained
Aleksa Gordić - The AI Epiphany
Non-Parametric Transformers | Paper explained
Aleksa Gordić - The AI Epiphany
Chip Placement with Deep Reinforcement Learning | Paper Explained
Aleksa Gordić - The AI Epiphany
Text Style Brush - Transfer of text aesthetics from a single example | Paper Explained
Aleksa Gordić - The AI Epiphany
Graphormer - Do Transformers Really Perform Bad for Graph Representation? | Paper Explained
Aleksa Gordić - The AI Epiphany
GANs N' Roses: Stable, Controllable, Diverse Image to Image Translation | Paper Explained
Aleksa Gordić - The AI Epiphany
VQ-VAEs: Neural Discrete Representation Learning | Paper + PyTorch Code Explained
Aleksa Gordić - The AI Epiphany
VQ-GAN: Taming Transformers for High-Resolution Image Synthesis | Paper Explained
Aleksa Gordić - The AI Epiphany
Multimodal Few-Shot Learning with Frozen Language Models | Paper Explained
Aleksa Gordić - The AI Epiphany
Focal Transformer: Focal Self-attention for Local-Global Interactions in Vision Transformers
Aleksa Gordić - The AI Epiphany
AudioCLIP: Extending CLIP to Image, Text and Audio | Paper Explained
Aleksa Gordić - The AI Epiphany
RMA: Rapid Motor Adaptation for Legged Robots | Paper Explained
Aleksa Gordić - The AI Epiphany
DALL-E: Zero-Shot Text-to-Image Generation | Paper Explained
Aleksa Gordić - The AI Epiphany
DETR: End-to-End Object Detection with Transformers | Paper Explained
Aleksa Gordić - The AI Epiphany
DINO: Emerging Properties in Self-Supervised Vision Transformers | Paper Explained!
Aleksa Gordić - The AI Epiphany
DeepMind DetCon: Efficient Visual Pretraining with Contrastive Detection | Paper Explained
Aleksa Gordić - The AI Epiphany
Do Vision Transformers See Like Convolutional Neural Networks? | Paper Explained
Aleksa Gordić - The AI Epiphany
Fastformer: Additive Attention Can Be All You Need | Paper Explained
Aleksa Gordić - The AI Epiphany
More on: Fine-tuning LLMs
View skill →Related Reads
📰
📰
📰
📰
Building the Next Generation of Synthetic Media – OpenCV Live Ep. 218
OpenCV Blog
Why Diffusion Transformers (DiTs) Are Replacing U-Nets in Generative AI
Medium · AI
Why Diffusion Transformers (DiTs) Are Replacing U-Nets in Generative AI
Medium · Deep Learning
How I Built FoodVision Big: Teaching a Model to Tell 101 Dishes Apart
Medium · Machine Learning
Chapters (3)
01:34 Intro
1:34
03:21 Hyperstack GPUs (sponsored)
3:21
59:18 Instruct Pix2Pix
🎓
Tutor Explanation
DeepCamp AI