AI Engineering for Art - with comfyanonymous
Key Takeaways
This video teaches AI engineering for art using comfyanonymous tools and techniques
Full Transcript
hey everyone welcome to the laden space podcast this is alesio partner and CTO desable partners and I'm joined by my co-host swix founder of small AI hey everyone we are in the chroma Studio again but with our first ever Anonymous guest comfy Anonymous welcome yeah well hello I feel like that's your full name you just go by comfy right yeah well a lot of people just call me comfy even though even even when they know my real name hey hey comfy yeah Wix is the same like you know not a lot of people call you yeah yeah you have a professional name right that people know you buy and then you have a legal name yeah it's fine how do I phrase this like I think people who are in the no know that comfy is like the tool for image generation and now other multimodality stuff I would say that when I first got started with stable diffusion the star of the show was automatic 111 right and I I actually looked back in my notes from 2022 is like comfy was already getting started back then but it was kind of like the upand comer and like your main feature was a flowchart can you just kind of rewind to that moment that that year and like you know how how you looked at the landscape there and decided to start comfy yeah I discovered stable diffusion in 2022 in October 2022 and uh well I kind of started playing around with it yes I and back then I was using automatic which was what everyone was using back then and I so I started with that cuz I had the CU When I started I had no idea like how the fusion models work how any of this work so oh yeah what was your prior background as an engineer uh just a software engineer yeah boring software engineer but like any any image stuff any orchestration distributed systems gpus no I was doing basically nothing interesting crud web development yeah well not web development just yeah some basic maybe some basic like automation stuff and okay just yeah no like no big companies or anything yeah but like already some interest in automations probably a lot of python yeah yeah of course python but like I wasn't actually used to like uh the node graph interface before I started confy UI it was just uh I just thought it was like oh like what's the best way to represent the diffus process in the user interface and I'm like oh W like natural oh this is the best way I found and this was a like with the node interface so how I got started was uh yeah so basically October 2022 just uh like I hadn't written a line of pie torch before that so it's a completely new what happened was I kind of got addicted to generating images as we all yeah and then I started experimenting with uh like the highrisk fixed in Auto which was uh for those that don't know the highr fixes just to gener since the diffusion models back then could only generate that low resolution so what you would do you would generate low resolution image then upscale then and of refine it again and that was kind of the hack to generate high resolution images I really liked generating like higher resolution images so I was experimenting with that and so I modified the code a bit okay what happens if I if I use different Samplers on the second fast so I was edited the code of Auto so what happens if I use a different sampler what happens if I use a different uh like different settings different number of steps and because back then the highrisk fix was very basic just so yeah now there's a whole library of just uh the up Samplers I think I think they added a bunch of uh of options to the highr fix since uh since since then but before that was just so basic so I wanted to go further I wanted to try okay what happens if I use a different model for the second second pass and then well then the auto code this was wasn't good enough for like it would have been uh harder to implement that in the auto interface then to create my own interface so that's when I decided to create my own and you were doing that mostly on your own when you started or did you already have kind of like a subgroup of people no I was uh on my own cuz it was just me experimenting with stuff so yeah that was it Bas so I started writing the code the January 1 2023 and then I released the first version on GitHub January 16 2023 that's how things got started and what what the name comy UI right away or yeah comy UI the reason the name my name is comfy is people thought my pictures were comfy so I just just name it it's my K she UI so yeah that's uh is there a particular segment of the community that you targeted as users like more intensive workflow artists you know compared to the automatic crowd or you know this was my way of like experimenting with uh with new things and like the highis fixed thing I mentioned which was like in comfy the first thing you could easily do was just chain different models together and then one of the first things I think the first times it got a bit of popularity was when I started experimenting with the different like applying prompts to different areas of the image yeah I called it area conditioning posted it it on Reddit and it got a bunch of up votes so I think that's when like when people first learned of comy UI m is that mostly like fixing hands oh no no that was just uh like let's say well it was very well it still is kind of difficult to like let's say you want a mountain you have an image and then okay I want a mountain here and I want the like a a fox here yeah so compositing the image Yeah by way was very easy it was just like a when you run the diffusion process you kind of generate okay you do pass one pass through the diffusion mod every step you do one pass okay this place of the image with this Prim this place place of the image with the other prompt and then the entire image with another prompt and then just average everything together every step and that was a area composition which I called it and then then a month later there was a paper that came out called multi- diffusion which was the same thing but yeah that's uh could you do area composition with different models or because you're averaging out you kind of need the same model you could do it with but yeah I had implemented it for different models but you you can do it with uh with different models if you want as long as the models share the same latent space like we we're supposed to ring a bell every time someone says lat space yeah like for example you couldn't use like EXL and SD 1.5 cuz those have a different Laten space but like uh yeah like SD 1.5 models different ones you could you could do that there's some models that try to work in pixel space right yeah they're very slow that's the problem that that's the the reason why stable diffusion actually became like popular like because was because of the latent space small in yeah because it used to be lat diffusion models and then they trained it up yeah cuz the pixel pixel diffusion models are just too slow so yeah have you ever tried to talk to like like stability the lat diffus guys like you know Robin Rach that crew yeah well I used to work at stability oh I actually know yeah I used to work at stability I got uh I got hired uh in June 2023 Ah that's the part of the story I didn't know about okay okay so the the reason I was hired is because they were doing sdxl at the time and they were basically sdxl I don't know if you remember it was a base model and then a refiner model basically they wanted to experiment like chaining them together and then uh they saw oh right come oh this we can use this to do that well let's hire that guy but they didn't they didn't pursue it for like sd3 what do you mean like the sdxl approach yeah the reason for for that approach was because basically they had two models and then they wanted to publish both of them so they they trained one on Lower time steps which was the refiner model and then they the first one while was trained normally and then they went during their test they realize oh like if we string these models together our like quality increases so let's publish that it worked yeah but like right now I don't think many people actually use the refiner anymore even though it is actually a full diffusion model like you can use it on its own and it's going to generate images I don't think anyone people have mostly forgotten about it but uh can we talk about models a little bit so stle diffusion obviously is the most known I know flux has gotten a lot of traction are there any underrated models that people should use more or what's the state of the union well the the latest uh State ofth art at least yeah for images there's yeah there's flux there's also SD 3.5 SD 3.5 is two models that there's a there's a small one 2.5b and there's the bigger one 8B so it's it's smaller than flux so and it's more uh creative in a way but flux yeah flux is the best people should give SD 3.5 a try cuz it's uh it's different I won't say it's better well it's better for some like specific use cases like if you want some to make something more like creative maybe as the 3.5 if you want to make something more consistent and flux is probably better do you ever consider supporting the closed Source model apis uh well they we do support them with as custom nodes we actually have some uh official custom nodes from uh different yeah I guess Dolly would have one yeah that's uh it's just not I'm not the person that handles that so sure sure quick question on on SD there's a lot of community discussion about the transition from SD 1.5 to sd2 and then sd2 to sd3 people still like you know very loyal to the previous generations of SDS uh yeah SD 1.5 and still has a lot of a lot of users the last based model yeah then sd2 was mostly ignored because it wasn't uh it wasn't a big enough improvement over the previous one okay so SD 1.5 sd3 flux and whatever XL XL that's the main one stable Cascade no stable Cascade that was a good model but uh that's the problem with that one is it got like sd3 was announced one week after yeah it was like a weird release uh what was it like inside of stability actually I mean statue limitations expired you know management has has moved so it's easier to talk about now yeah the inside still ability actually that model was ready uh like three months before but it got stuck in uh red teaming so basically the proba if that M had released it was supposed to be released by the authors then it would probably have gotten very popular since it's a it's a step up from sdxl but it got all of its momentum stolen by the sd3 announcement so people kind of didn't develop anything on top of it and so it's yeah it was a good model at least uh completely mostly ignored for some reason like it see I think the naming as well matters it seemed like a branch off of the main main tree of developments yeah well it was different researchers that did it like yeah very like a good model like it's the worsters shin authors I not know if I'm pronouncing it worin yeah yeah actually met them in uh Vienna yeah they worked at stability for a bit and they left right after the Cascade release this is Dustin right no Dustin's sd3 no Dustin is sd3 sdxl that's Pabo and do I think I'm pronouncing his name correctly yeah that's ve very good it seems like the community is very they move very quickly yeah like when there's a new model out they just drop whatever the one is and they just all move wholesale over like they don't really stay to explore the full capabilities like if if the stable Cascade was that good they would have AB tested a bit more instead they're like okay sd3 is out like let's go you know well I find the opposite actually the community doesn't like they only jump on a new model when there's a significant Improvement I see like if there's a only like a incremental Improvement which is what uh most of these models are going to have especially if you do stay the same parameter count yeah like you're not going to get a massive Improvement into like unless there's something big that that changes so yeah and how are they evaluating these improvements like um because there it's a a whole chain of you know comfy workflows yeah how does how does one part of the chain actually affect the whole process are you talking on the model side specific model specific right but like once you have your whole workflow based on on a Model it's very hard to move uh not well not really well depends on your depends on the specific kind the workflow yeah so I do a lot of like text and image yeah it's when you do change like most workflows are kind of going to be compatible between different models it's just like you might have to completely change your prompt completely change okay well I mean that maybe the question is really about evals like what does the comfy community do for evals just you know well that they don't really do it's more like oh I think this image is nice that's they just subscribe to fur Ai and just see like you know what fer is doing yeah they just they just generate like like I don't see anyone really doing like at least on the comfy side comfy users they it's more like oh generate images and see oh this one's nice this like yeah yeah it's not uh like the the more like scientific like like checking that's more specifically on like model side if uh yeah but there is a lot of Vibes also because it is a like artistic uh you can't create a very good model that doesn't generate nice images CU most the images on the internet are ugly so if you if that's like if you just oh I have the best model that can like it's super smart I train it on all the like I train on just all the images on the internet the images are not going to look good so yeah yeah they're going to be very consistent but yeah people like it's not going to be like the the look that people are going to be expecting from uh from a model so can we talk about loras because we talk we talk about models then like the next step is probably Laura's before actually I'm kind of curious how Laura's entered the tool set of the image Community because the Laura paper was 2021 and then like there was like other methods like textual inversion that was popular at the early SD stage yeah I can't explain the difference between textural inversions that's basically what you're doing is you're you're training a cuz well yeah stable diffusion you have the diffusion model you have text encoder so basically what you're doing is training a vector that you're going to pass to the text decoder it's basically you're training a new word yeah it's a little bit like representation engineering now yeah yeah basically yeah you're just uh so yeah if you know how like the text and coder Works basically you have a you take your your words of your prompt you convert those into tokens with the tokenizer and those are converted into vectors basically yeah h token represents a different Vector so each word presents a vector and those depending on your words that's the list of vectors that get passed to the text encoder which is just uh yeah just a stack of of attention like basically it's a very close to llm architecture Yeah Yeah so basically what you're doing is just training a new Vector we're saying well I have all these images and I want to know which word does that represent and it's going to get like you train this vector and then and then when you use this Vector it hopefully generates uh like something similar to your images yeah I would say it's like surprisingly sample efficient in picking up the concept that you're trying to train it on yeah well people have kind of stopped doing that even though back at like when I was at stability we we actually did train internally some like Tex versions on like t5x XL actually work pretty well but for some reason yeah people don't use them and also they might also work like like yeah that's is something it probably have to test but maybe if you train a textural inversion like on t5x XL it might also work with all the other models that use t5x XL because same thing with like U like the texal inversions that uh that were trained for SD 1.5 they also kind of work on sdxl because sdxl has the has two text encoders and one of them is the same as the as the SD 1.5 clip l so those they actually they don't work as strongly because they're only apply to one of the text encoders but and the same thing for sd3 three sd3 has three text encoders so it works it still you can still use your text from version SD 1.5 on sd3 but it's just a lot weaker because now there's three text encoders so it gets even more diluted yeah do people experiment a lot on just on the clip side uh there's like sigp there's blip like do people experiment a lot on on you can't really replace yeah cuz they're trained together right yeah they're trained together so you can't uh like what I've seen people experimenting with is a long clip so basically someone fine-tune the clip model to accept longer promise oh oh it's kind of like long context fine tuning yeah so so like it's it's actually suppored in core comy how long is long regular clip is 77 tokens yeah one clip is 256 okay so but the hack that uh like you if you use table diffusion 1.5 you've probably noticed oh it still works if I if I use long prompts promps longer than 77 words well that's because the hack is to just uh well you split you split it up in chugs of 7even your whole your big prompt let's say you you give it like the massive text like the Bible or something and it would split it up in try seven and then just pass each one through the clip and then just everything together at the end it's not ideal but it actually works like the positioning of the words really really matters then right like this is why order matters in prompts yeah yeah like it it works but it's it's not ideal but it's what people expect like if if someone gives a huge prom they expect at least some of the concepts at the end to be like present in the image but usually when they give long prompts they they don't like they don't expect uh like detail I think so that's why it works very well and while we're on this topic uh prompt waiting NE negative prompting all all sort similar part of this layer of the stack yeah the the hack for that which works on clip like it basically it's just uh for SD 1.5 well for SD 1.5 The Prompt plating works well because the clip L is a is not a very deep model so you have a very high correlation between you have the input token the index of the input token vector and the output token they're very the concepts are very close closely l so that means if you interpolate the vector from what well the the way comi does it is it has okay you have the a vector you have a empty prompt so you have a a chunk like a clip output for the NP promt and then you have the one for your prompt and then it interpolates from that depending on your prompt weight the weight of your of your tokens so so if you yeah so that's how how it does a promptu waiting but this stops working the deeper your text encoder is so on T5 XSL it doesn't work at all so wow is that a problem for people I mean because I'm used to just move moving up numbers not uh well so you just use words to describe right cuz it's a bigger language model yeah yeah so honestly it might be good but I haven't seen many complaints on flux outy it's not working so cuz I guess people can sort of get around it with the with language so yeah and then coming back to loras now the the popular way to to customize models is loras and I I saw you also support lcon and loha which I've never heard of before there's a bunch of cuz oh what what the doora is essentially is uh instead of uh like okay you have your your model and then you want to find tune so instead of uh like what you could do is you could fine tune the entire thing tune but that's a bit uh heavy so to speed things up and make things less heavy what you can do is just fine tune some smaller weights like basically two two matrices that when you multiply like two low rank matrices and when you multiply them together gives a represents a difference between trained weights and your base weights so by training those two smaller matrices that's a lot less thing yeah and they're portable so you can share them it's easier smaller yeah that's the how luras work so basic so when when inferencing you can inference with them pretty efficiently like how why does it it just when you use Aur it just applies it straight on the weights so that there's only a small delay at the B like before the sampling to when it applies the weights and then it just same speed as uh as before so for for inference it's it's not that bad but uh and then you have so basically all the Lura types like loha loan everything that's just different ways of representing that like basically you can call it kind of like compression even though it's not really compression it's just different ways of representing like just okay I want to train a different on the difference on the weights what's the best way to represent that difference there's the basic Laura which is just oh let's multiply these two matrices together and then there's all the other ones which are all different algorithms so so let's talk about what confy UI actually is I think most people have heard of it some people might have seen screenshots I think fewer people to build very complex workflow so when you started automatic was like the super simple way what were some of the choices that you made so the node workflow is there anything else that stands out as like this was like a unique take on how to do image Generation workflows Well I feel like yeah back there everyone was tried to make like easy to use interface I'm like well everyone's trying to make an easy to use interface let's make a hard to use interface like so like I like I don't need to do that everyone else doing it so let me try something uh like let me try to make a powerful interface yeah that's not easy to use so so like yeah there's a sort of node execution engine your read me actually list has really good list of features of things you prioritize right like um let me see like uh sort of re-executing from from any parts of the workflow that was changed asynchronous Q system smart memory management like all this seems like a lot of engineering that yeah there's a lot of Engineering in the in the back end to make things as I was always focused on making things work locally very well cuz that's CU I was using it locally so everything uh so there's a lot of uh a lot of thought and work and like getting everything to run as well as possible so yeah Ki is actually more of a back end at least well now now the front end's getting a lot more development but but before before it was I was pretty much only focused on the back end yeah so v0.1 was only August this year yeah before version name so yeah and so what was the big rewrite for the 0.1 and then the 1.0 uh well that's more on the front end side that's cuz before that it was just uh like the UI what cuz when I first wrote it I just uh I said okay how can I make like I can do web development but I don't like doing it like what's the easiest way I can slap a node interface on this and then I found this Library light graph like JavaScript library life graph light graph usually people will go for like react flow for like a flow build yeah but that seems like too complicated so I didn't really want to spend time like developing the the front end so I'm like well oh light graph this has the whole no interface so okay let me just plug that into to my back end then I feel like if streamlet or gradio offered something you would have use streamlet or gradio cuz it's Python streamlet and gr like gradio I don't like gradio it's like the that's that's one of the reasons why like automatic was very bad it's great because uh the problem with gradate it forces you to well not forces you but it kind of uh makes your your interface logic and your backend logic and that just sticks them together it's supposed to be easy for you guys for if you're a python main you know I'm a JS main right if you're a python main it's supposed to be easy yeah it's well it's but it makes your whole software a huge mess I see I see so you're mixing concerns instead of separating concerns well it's cuz like front end and back end front end and back end should be well separated with a defined API like that's that's how we supposed to do it smart people disagree just stick sticks everything together it makes makes it easy to like make a huge mess and also it's the that there's a lot of issues with uh with radio like it it's very good if all you want to do is just get like slap a quick interface on your uh like to to show off your like your ml project like that's what it's made for yeah yeah like like there's no problem using it like oh I have my I have my code I just wanted in quick interface on it that's perfect like use gray but if you want to make something that's like a real like real software that will last a long time and will be easy to maintain then I would avoid it yeah so so your criticism is streamlit and gradio the same I mean those are the same criticisms yeah stre I haven't haven't used as much yeah I just looked a bit similar philosophy yeah it's similar it's just it just seems to me like okay for quick like AI demos it's perfect yeah going back to like the the the core Tech like asynchronous cues slow re-execution smart Mery management you know anything that you you you're very proud of or was very hard to figure out yeah the thing that's the biggest pain in the ass is probably the memory management yeah were you just paging models in and out or yeah before it was just okay load the model completely unload it load the new model completely unloaded then okay that that works well when your model are small but if your models are big and it takes like it's and someone has a like a a490 and the model size is 10 gigabytes that can take a few seconds to like load and load load and load so you want to try to keep things like in memory in the GPU memory as much as possible what kyui does right now is that uh it tries to like estimate okay like okay you're going to sample this model it's going to take probably this amount of memory let's remove the models like this amount of memory you Lo that's been loaded on the GPU and then just execute it but so there's a fine line between just cuz try to remove the least amount of models that are already loaded cuz I SP like Windows driver and one another problem is uh the Nvidia driver on Windows by default because there's a way to there's an option to disable that feature but by default it um like if you start loading you can overflow your GPU memory and then it's the driver is going to automatically start paging to Ram but the problem with that is it it makes everything extremely slow so when you see people complaining oh this model it works but oh [ __ ] it starts slowing down a lot that's probably what's happening so it's basically you have to just try to get use as much memory as possible but not too much or else things start slowing down or people get out of memory and then just find try to find that line where oh like the driver on window starts paging and stuff yeah and yeah and the problem with pie torch is it's uh it's high levels don't have that much fine grain control over like specific memory stuff so kind of have to leave like the memory freeing to to Python and Pie torch which is can be annoying sometimes so you know I think one thing is a as a maintainer of this projects like you're designing for a very wide yeah surface area of compute like you even support CPUs yeah well that's that's just F torch F torch so yeah it's just that's not that's not hard to support first of all is there a market share estimate like is it like 70% Nvidia and like 30% AMD and then like miscellaneous on Apple silicon or whatever uh for comy yeah yeah yeah I don't know the market share can you guess uh I think it's mostly Nidia yeah because am the problem like AMD Works horribly on Windows like Linux it it works fine it's it's slower than the price equivalent Nidia GPU but it works like you can use it generate images everything works on Linux on Windows you might have a hard time so that's the problem and most people I think most people who bought the AMD probably use Windows they probably aren't going to switch to Linux so so until AMD actually like ports their like raw cam to to Windows properly and then there's actually P torch I think they're they're doing that they're in the process of doing that but uh until they get they get a good like pie torch Rock and build that works on Windows it's uh like they're going to have a hard time yeah we got to get George on it yeah well he's trying to get Lisa Su to do it but let's talk a bit about like the node design so unlike all the other text to image you have a very like deep so you have like a separate node for like clip and code you have a separate note for like the cas sampler you have like all these notes going back to like the making it easy versus making it hard but like how much do people actually play with all the settings you know kind of like how do you guide people to like hey this is actually going to be very impactful versus this is maybe like less impactful but we still want to expose it to you uh well I try to expose uh like I try to expose everything or but so yeah at least for the but for things like for example for the Samplers like there's uh like yeah four different sampler nodes which go in easiest to most advanced so yeah if you go like the easy node the regular sampler node that's you have just the basic settings but if you use like the sampler Advan custom Advanced node that that when you can actually you'll see you have like different nodes I'm looking it up now yeah what are like the most impact impactful parameters that you use so it's like you know you can have more but like which ones like really make a difference yeah they all do they all have their own like they all like for example yeah steps usually you want steps you want them to be as low as possible but you want if you're optimizing your your workflow you want to you lower the steps until like the images start deteriorating too much cuz that um yeah that's the number of steps you're you're running the diffusion process so if you want things to be fast that's so lower is better but yeah CFG that's more you can kind of see that as the contrast of the image like if your image looks too burnt out then you can wear the CFG so yeah CFG that's how yeah that's how strongly the like the negative versus positive prompt so when you sample a diffusion model it's it's basically a negative prompt it's just yeah positive prediction minus negative prediction contrastive loss yeah just positive minus negative and the CFG that's the multiplier yeah yeah so what are like good resources to understand what the parameters do I think most people start with automatic and then they move over and it's like step CFG sampler name scheduler D noise read it but honestly well it's more it's something you should like like try out yourself I don't you you don't necessarily need to know how it works to like what it does cuz even if you know like CFG it's like positive minus negative Pro yeah so the only thing you know with CFG is if it's 1.0 then that means the negative prpt isn't applied also mean sampling is two times faster but uh yeah but other than that it's more like you should really just see what it does to the images your yourself and you'll you'll probably get the more intuitive understanding of what these things do mhm any other noes or things you want to shut out like I know the anime D IP adapter those are like some of the most popular ones um yeah what else comes to mind uh not notes but there's uh like what I like is when when some people sometimes they make they make things that they use can Pui as their back in like there's a a plugin that for create that uh uses confi as its back end so you can use like all the models at work in comfi in CR and I think I've tried it once but I know a lot of people use it and probably do nice so what's the craziest node that people have built like the most complicated craziest no like I like yeah I know some people have uh made like video games and in comfy with like stuff like that so like someone like I remember like yeah last I think it was La last year someone made a like a like Wolfenstein com and then one of the inputs was oh you can generate the texture and then it changes the texture in the game so you could plug it to like workflow and there's a lot of if you look there there's a lot of crazy things people do so yeah and now there's like a node register that people can use to like download nodes and yeah like well there's always been like the confi manager but we're trying to make this more like I know official like uh with yeah with the the node registry because before before the node registry the it like okay how did your custom node get in compy manager that's the guy running it who like every day he searched GitHub for new custom nodes and added them manually to his to his custom node manager so we're trying to make it uh less effort for him basically yeah but I was looking I mean there's like a YouTube download node there's like this is almost like you know a data pipeline more than like an image generation thing at this point it's like you can get data in you can like apply filters to it you can generate data out yeah you can do a lot of uh different things yeah something I think uh what I did is uh I made it uh easy to make custom nodes so I think that that helped a lot for the the ecosystem because it is very easy to just make a node so yeah a bit too easy sometimes then then we have the issue where there's a lot of custom note packs which share similar notes so but uh well that's uh yeah something we're trying to solve by maybe bringing some of the functionality into core yeah and then there's like video people can do video generation yeah video that's well the the first video model was like stable video diffusion which was last yeah exactly last year I think like when year ago but that wasn't the true video model so it was like it was like moving images yeah I generated video what I mean by that is it's like it's still 2D Laten it's basically what they did is they took sd2 and then they added some temporal attention to it and then trained it on videos and know so it's it's kind of like animated diff like close same same idea basically why I say it's not a true video models you still have like the 2D latens like a true video model like Mochi for example would have 3D latens so you can like move through the space basically it's the the difference you're not just kind of like reorienting yeah and it's also well it's also because you have a temporal V8 that also like Mochi has a temporal vae that compresses on like the temporal direction also so that's something you don't have with like yeah animate diff and the stable video diffusion they only like press spatially not temporally so yeah so these models that's why I call them like true video models there there's yeah there's actually a few of them but uh the the one I've implemented in comfy is moochi because that that seems to be the best one so far yeah we had AJ come and speak at the state of diffusion Meetup other open one I think I've seen is coog video Yeah C video yeah that one see yeah it also seems decent but uh yeah but Chinese so we don't use it's fine it's just yeah I could yeah it's just it uh there's it's not the only one there's also a few others which I the rest are like closed stores right like cling yeah the closed stores there's a bunch of them but I mean open I've seen a few of them like I can't remember their names but there's Cog Cog videos the big the big one then there's also a few of them that released at the same time yeah there's one that released at at the same same time SSD 3.5 same day which is why I don't remember the name we should have a release schedule so we don't conflict on each of these things yeah I think SD 3.5 and mochi released on the same day so everything else was kind of drown completely drowned out so uh for some reason lots of people pick that they to release their stuff yeah which is well shame for those in gas and think Omni Jen also we always the same day which also seems interesting but yeah yeah what's comfy so you are comy and then there's like com.org um I know we do a lot of things for like news research and those guys also have kind of like a more open source and on uh thing going on how do you work like you mentioned you mostly work on like the the core piece of it and then what maybe I should f it because yeah I yeah I feel like maybe yeah I only explained part of the story right yeah yeah maybe I should explain the rest so yeah so yeah basically uh January that's when the first January 2023 January 16 2023 that's when M was first released to the public then yeah did a Reddit post about the area composition thing somewhere in uh I don't remember exactly May the end of January beginning February and then some a YouTuber made a video about it uh like Olivio he made a video about comfy in March 2023 I think that's when it first real burst of attention and by yeah that time I was continue to developing it and it was getting uh people were starting to use it more which unfortunately meant that I my as I had first written it to do like experiments but then well my time to do experiments when started going down cuz yeah cuz yeah people were actually starting to use it and like I had to and I said well yeah time to add all all these features and stuff yeah and then I got hired by stability June 2023 then I made the basically yeah they hired me because they wanted the do sdxl so I got sdxl working very well thei because they were experimenting with it actually the sdx how the sdxl released worked is they released for some reason like they released the code first but they didn't release the model checkpoint oh yeah so they released the code and then well if since the research was relas the code I released the code in C 2 and then the checkpoints were basically Early Access people had to sign up and they only allowed the a lot of people from Ed do emails like Ed do email like like they gave you access basically to the zero sdxl 0.9 and well that uh leaked right of course because of course it's going to leak if you do that H well the only way people could easily use it was with comfy so yeah people started using it and then I fixed the few all the issue people had so then the big 1.0 release happened that and well Ki was the only way a lot of people could actually RN it on their computers cuz it just like automatic was so like inefficient and bad that most people couldn't act like it just wouldn't wouldn't work like cuz he he did a quick implementation so people were forced to use kyui and that's how it became popular because people had no choice the growth hack yeah yeah yeah like everyone like people people who didn't have the 49 they had like who had just regular gpus yeah yeah they just they didn't have a choice so yeah I got a 47 so think of me and so today what's is there like a core comy team or uh yeah well right now um yeah we are hiring actually so right now cor could like on the core core itself it's it's me and but because uh reason where F like all the focus has been mostly on the front end right now cuz that's the thing that's been neglected for a long time so uh so most of the Focus right now is uh all on the front end but we are yeah we will soon get more people to like help me with the actual backend stuff because that's once the once we have our V1 release which is going to be packaged confused why with the nice interface and easy to install on Windows and hopefully Mac yeah yeah once we have that uh we're going to have to lots of stuff to do on the back inside and also the front side but uh what's the release I'm on the waight list What's the timing uh soon soon yeah like I don't want to promise I release because yeah we we do have a a real estate weth targeting but yeah I'm not sure if if it's public yeah yeah and how we're going to like we're we're still going to continue like doing the open source like makingi the best way to run uh like stable diffusion models like at least the the open source side and like it's going to be best way to run or models locally but we will have a few like a few things to to make money from it like uh Cloud inference or like that type of that type of thing so and maybe some like some things for some Enterprises I mean a few questions on that how do you feel about the other comfy startups I mean I think it's great they're using your name you know yeah well it's better to use comfy than to use something else yeah that's true yeah like yeah it's fine I don't like we're like yeah I'm we're going to try not to we don't want to like we want them to people to use comfy CU like I said it's better that people use comfy than something else so as long as they use comfy it's a I think it helps it helps the ecosystem because more people even if they don't like even if they don't contribute directly the fact that they are using comfy means that like people are more likely to like join the ecosystem so yeah and then would you ever do text yeah well that you can already do text with some custom notes so yeah it's something we we like yeah it's something I've wanted to eventually add to core but it's more like not not a very high priority but uh because a lot of people use text for like prompt enhancement and like other things like that so it's uh yeah it's just that my focus has always been like on the US models yeah unless some tax diffusion model comes out yeah David Holtz is investing a lot in tax diffusion well if if a good one comes out then well I'll probably implement it since it fits with the whole yeah I mean I I imagine it's going to be close source to M Journey so yeah well yeah if an open one comes out yeah then uh yeah I'll probably yeah yeah I'll probably implement it just no cool confy thanks so much for for coming on this is fun [Music]
Original Description
Full show notes: https://www.latent.space/p/comfyui
Happy new year friends! Thanks for all the love on the Latent Space Live and 100th Episode End of Year recap. Your support has boosted us 30 places in the Podcast charts, and that always helps us book great guests and organize more industry events for you! We don't say this enough but thank you to everyone who has left a review on Apple Podcasts or subscribed to our new YouTube channel.
Last year we broke new ground when we interviewed our first public company CEO with Drew Houston, and first technology Cabinet member with Minister Josephine Teo, and first year with full coverage of leading labs across Meta, OpenAI, Anthropic, Reka, and Google DeepMind. For our 101st episode, we are proud to introduce another first, with our first anonymous guest!
As swyx mentions in the episode, Latent Space was started in the immediate aftermath of Stable Diffusion, and the uncredentialed software engineers it enabled set the stage for the LLM wave that was to come with ChatGPT. The earliest winner of the Stable Diffusion tooling wars was SD Web UI, a Gradio app by the anonymous young creator Automatic 11 11 that quickly amassed over 100,000 GitHub stars for how it rapidly shipped plugins and usable interfaces for the rapidly growing Stable Diffusion ecosystem.
However, these days, the power tool of choice is now ComfyUI, by today's guest, Comfy anonymous, who is gracing us with his first ever podcast appearance today. The shift from Automatic 11 11 to ComfyUI reflects a shift away in the image diffusion space from prompting and tweaking settings in 2022, to more complex and parallel workflows chaining together different models, and orchestrating long running operations that can also include video processing, visualized on an intuitive canvas instead of long YAML or code blocks. Because ComfyUI is open source, there are now multiple YCombinator startups built off of a Comfy workflow, or offering ComfyUI as a service directly.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Latent Space · Latent Space · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Ep 18: Petaflops to the People — with George Hotz of tinycorp
Latent Space
FlashAttention-2: Making Transformers 800% faster AND exact
Latent Space
RWKV: Reinventing RNNs for the Transformer Era
Latent Space
Generating your AI Media Empire - with Youssef Rizk of Wondercraft.ai
Latent Space
RAG is a hack - with Jerry Liu of LlamaIndex
Latent Space
The End of Finetuning — with Jeremy Howard of Fast.ai
Latent Space
Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
Latent Space
Powering your Copilot for Data - with Artem Keydunov from Cube.dev
Latent Space
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Latent Space
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Latent Space
The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Latent Space
The AI-First Graphics Editor - with Suhail Doshi of Playground AI
Latent Space
The Accidental AI Canvas - with Steve Ruiz of tldraw
Latent Space
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Latent Space
The Four Wars of the AI Stack - Dec 2023 Recap
Latent Space
The State of AI in production — with David Hsu of Retool
Latent Space
Building an open AI company - with Ce and Vipul of Together AI
Latent Space
Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
Latent Space
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
Latent Space
Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Latent Space
Making Transformers Sing - with Mikey Shulman of Suno
Latent Space
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
Latent Space
Why Google failed to make GPT-3 -- with David Luan of Adept
Latent Space
Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Latent Space
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
Latent Space
Breaking down the OG GPT Paper by Alec Radford
Latent Space
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Latent Space
This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Latent Space
LLM Asia Paper Club Survey Round
Latent Space
How to train a Million Context LLM — with Mark Huang of Gradient.ai
Latent Space
How AI is Eating Finance - with Mike Conover of Brightwave
Latent Space
How To Hire AI Engineers (ft. James Brady and Adam Wiggins of Elicit)
Latent Space
State of the Art: Training 70B LLMs on 10,000 H100 clusters
Latent Space
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Latent Space
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Latent Space
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Latent Space
Synthetic data + tool use for LLM improvements 🦙
Latent Space
RLHF vs SFT to break out of local maxima 📈
Latent Space
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
Latent Space
Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
Latent Space
Answer.ai & AI Magic with Jeremy Howard
Latent Space
Is finetuning GPT4o worth it?
Latent Space
Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Latent Space
Building AGI with OpenAI's Structured Outputs API
Latent Space
Q* for model distillation 🍓
Latent Space
Finetuning LoRAs on BILLIONS of tokens 🤖
Latent Space
Cursor UX team is CRACKED 💻
Latent Space
Choosing the BEST OpenAI model 🏆
Latent Space
How will OpenAI voice mode change API design?
Latent Space
STEALING OpenAI models data 🥷
Latent Space
[Paper Club] 🍓 On Reasoning: Q-STaR and Friends!
Latent Space
[Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Latent Space
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Latent Space
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
Latent Space
Prompt Engineer is NOT a job 📝
Latent Space
Prompt Mining LLMs for better prompts ⛏️
Latent Space
The six pillars of few-shot prompting 🔧
Latent Space
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Latent Space
[Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Latent Space
Can you separate intelligence and knowledge?
Latent Space
Related Reads
📰
📰
📰
📰
Use a model route manifest before Dify, Cursor, and Node.js share Vector Engine
Dev.to AI
What 90 Days of Comments on AI Side Panels Taught Me About Distribution
Dev.to · AI Buddy
Claude vs ChatGPT for Small Business (2026): The Real Numbers Behind the Quiet Takeover
Medium · AI
AI Builds the Website in Minutes. So What Are You Actually Charging For?
Medium · AI
🎓
Tutor Explanation
DeepCamp AI