ComfyUI Tutorial Series: Ep09 - How to Use SDXL ControlNet Union

pixaroma · Beginner ·🎨 Image & Video AI ·1y ago

Key Takeaways

This video teaches how to use SDXL ControlNet Union in ComfyUI

Full Transcript

welcome to episode nine of our comfyi tutorial Series today I want to talk about control net specifically the union control net for sdxl uh this model helps the AI create images that match a specific composition I will show you how to integrate it into your workflow and provide a few examples I will add a link to this hugging face page in the description till now we had separate models for specific uses like pose or canny uh requiring multiple models however now we have a unified model model that includes all these options in one this model contains 12 controls you can find more information and examples on the page to download the model go to the files in the version section you'll see two models listed the one we want is the new one called the Promax click the download button then go to comfy UI navigate to the models folder look for control net and save the downloaded file in this folder so we've downloaded this model but what does it do control net is a guiding tool for stable def Fusion that helps the eye create images according to a specific structure or style you want and think of it like an artist who skilled at painting but sometimes needs a sketch to get started a control net acts like that initial sketch guiding where and how to paint uh training a control net model involves teaching it to match specific inputs with the desired images using a set of examples uh control net models are designed to work only with the particular AI model they were made for for instance a control net for SD 1.5 Works only with SD 1.5 models while a control net for sdxl is compatible only with sdxl back to comfy UI if you go to the manager and update your comfy UI you should get access to a new interface now when you double click on the canvas you can see a preview of each node on the left allowing you to check if it's the right node before adding it in the settings you can find more options for example you can enable or or disable certain menus and adjust their position to the top or bottom currently I have a bar at the bottom and a sidebar with a node Library here you can search for nodes expand folders and drag and drop nodes into the interface you also have a node map where you can activate or deactivate nodes by clicking on the little I icon at the bottom there's the manager and a q button from here you can toggle the theme the settings button allows you to change the theme color options include light solarized Arc obsidian and others I'll leave it on the default for now if you go back to settings and scroll down you can revert to the original interface by disabling the new menu option now back to control net sorry for the delay I just thought I should share this to work with the control net model we need some custom notes go to the manager then custom nodes manager and search for the term Venture install the comfy UI art Venture node which includes a control net pre-processor needed for our workflow next search for pre-processors and install the control net auxiliary pre-processors node to get access to additional pre-processors finally search for comfy roll and install comfy roll Studio to stack multiple control net models after installation click on restart it may take a while to install everything but once it's done you'll be back to comfy UI let's start simple and explain step by step use the clear button to get an empty canvas add your first node called control net pre-processor from art Venture you can see the preview of this node on the left this node has an input and output for an image so connect a load image note on the left and either a preview or save image node on the right if you run it and nothing happens it's because the pre-processor is set to none by the way if you're wondering how while I have the runtime displayed on top of the node it's due to a previous node I installed in another episode if I go to the manager and custom nodes manager and sort for installed nodes you can see one called comfy UI easy use in the pre-processor settings you can choose from a list list of options let's go with canny running it will give you a canny map what does this pre-processor do before the image goes through the AI model it's analyzed by the pre-processor it acts like a filter or scanner that extracts certain features from the image such as edges outlines or depth it simplifies the input image into something the AI can easily understand think of the pre-processor as creating a blueprint if you're building a house you don't start by randomly placing bricks you first make a blueprint that shows where everything will go the pre-processor creates this blueprint for the AI from here you can change the type of model you're using I will choose the sdxl version even though I haven't added the sdxl workflow yet just so you can get familiar with the settings for resolution you can choose either 512 or 1024 running it at a higher resolution provides more pixels and better detail you can go higher if you need to capture fine details but there is a limit going too high will result in an error if you upload another image like this one of a woman you'll see that the choice of pre-processor can affect the level of detail captured different pre-processors pick up different details so you may need to experiment depending on your needs you can also adjust the resolution to see if it captures more detail in the map switching to something like depth will produce a completely different map you can use multiple pre-processors on the same image for example if you duplicate the nodes you can connect them to get both a depth and cany map for an image with a building now let me quickly explain some of the pre-processors I use most often canny detects edges in an image think of it like tracing the outline of an object with a pencil it highlights the major lines and borders it's useful for or generating images where the structure of objects is important such as detailed line art or architectural designs depth understands the 3D structure of an image imagine looking at a photo and recognizing what is closer to you and what is farther away it helps create images with a sense of depth making it useful for realistic scenes where perspective matters mlsd uh detect straight lines in an image such as those in architectural photos it's like using a ruler to draw straight lines on a sketch highlighting the structure of buildings or other geometric forms it's great for architectural designs you do it cityscapes or any scene where straight lines are prominent uh normal map analyzes the surface normals which are vectors indicating the direction surfaces are facing and think of it as understanding the texture of of a surface you know such as whether it's uh bumpy flat or curved it's useful for generating detailed texture or enhancing the realism of 3D like images uh open pose detects human poses by identifying key points on the body like joints imagine drawing a stick figure to show how a person is standing or moving it's ideal for images where the position and the movement of people are important such as in action shots or character designs scribble allows users to draw simple sketches or scribbles that control net interprets as the structure for the final image it's like creating a rough draft and having an artist turn it into a polished piece of art it's great for quickly sketching out a concept and having the AI fill in the details segmentation breaks down an image into different regions or segments based on content like Sky ground or objects imagine coloring different areas of a picture with various colors uh to separate each element it's it's useful for scenes where different parts of the image need to be treated separately such as in complex environments or when creating layered compositions for some of the pre-processors I encountered an error but updating comfy UI after editing this video resolved it the issue was related to downloading a file always check the command window for information about what went wrong as it can provide details about the era that you can search for online let's integrate it into a workflow load and sdxl basic text to image workflow I'm similar to what we did in episode 3 as you can see now uses the Juggernaut X sdxl model it processes the prompts along with the empty latent image then goes to K sampler and is finally decoded to produce the final image make some space between the prompts and the case sampler by the way in the new version if you hover over nodes and Fields you'll get a brief description of what each one does let's integrate it into the workflow to make it easier to understand think of it like this we want to apply the control net to this workflow first search for the node called apply control net this comes with comfy UI so position it here now you'll see that the conditioning is connected directly to the positive we need to make a detour and connect it to apply control net instead connect the conditioning to conditioning and then conditioning to positive this node also needs a control net and an image drag a link from the control net and select load control net model which comes with comfyi in this note you can see the control net Union Pro Max model that we downloaded at the beginning of the video if you have more control net models they should appear in this drop-down menu for the image input we need a specific type of image that the AI can understand so add the node we use to test different pre-processors called control net pre-processor from art venture this node will help prepare the image in a way that control net can work with effectively I will select a pre-processor from the list for example canny then for the SD version I'll choose sdxl since we're using an sdxl model and control net for the value I'll go with 1024 move the node here underneath and then connect the image output to the image input on the apply control net node now select a load image node because we need an image for the pre-processor to work connect the output of the load image node to the image input of the control net pre-processor let's change the example image to a warrior image now the positive prompt goes to apply control net along with the control net model the image converted by the pre-processor into an image map that the AI can understand also goes to apply control net all these then go to the K sampler now I'd like to add a preview node to this node to see the resulting image and check if it captures the details I want and that's the complete workflow let's test it if the result isn't what you expected it might be because the prompt didn't accurately describe the image for instance if you add a prompt like a robot holding a baseball bat you can see that using canny keeps most of the lines and features of the warrior in this case one option is to try a different pre-processor for instance open pose could be a good alternative but I'll show you that later for now let's try using depth instead now it's looking better but the map is capturing the handle of the sword so it will be difficult to get a baseball bat if it's constrained by control net to look like a sword you can either switch to open pose or adjust your prompt to describe a robot holding a sword or something with a similar shape now you get this awesome image we started with a Warrior Man and ended up with a warrior robot by the way if I had used open pose with this prompt I would have gotten this image here here's how I approach it first I think about what I want to create and craft a prompt for it in this case I wanted a robot with a baseball bat next I search for an image that matches the pose I have in mind and I found a free image online of a baseball player then I experiment with different pre-processors in this scenario I didn't want all the lines and details that canny would provide and open pose wouldn't capture the bat effectively since I needed the bat to be in the right position dep was the best option as you can see the resulting image is quite cool back to the workflow I'm replacing the warrior image with a picture of a house I'll prompt for a futuristic building in Winter as you can see the depth pre-processor captures the details I need and the result is this nice image it's similar to the building I have but can be generated in any color style or season you can even modify it to be something else like a cake or anything with a similar shape mlsd is also effective with architecture since it extracts the straight lines from the image this is the result using mlsd here's how it looks with a normal map pre-processor for this example I changed the prompt to reflect an Autumn season let's load an image of an anime girl for the prompt I want a cute cartoon bunny I'll choose the scribble pre-processor as you can see it captures almost all the details from the illustration when you run it you get a variation of that image if I switch to canny I get a different type of variation because it interprets lines differently but since I've prompted for a bunny how can I get a bunny in the apply control net node there's a strength value when it's set to one control net strongly influences the results making them closely align with the input image if you reduce the strength value the variation becomes less influenced by the original image at a strength value of 0.3 three the bunny starts to take shape and at 0.2 we begin to see our bunny more clearly so even if you can't find the exact image for your control net you can still achieve your desired result by adjusting the strength value let's upload a different image like a portrait of a woman I'll prompt for a woman with blonde hair in the winter wearing a red blouse as you can see using canny didn't capture all the details in the face because there wasn't enough contrast between the shapes to accurately capture them this shows that the quality of the image is important the clearer and more contrasting the elements in the image the more details the pre-processor can extract the result is quite nice while the face is different the overall composition Remains the Same now let me show you how to combine or stack multiple control net models search for a node by typing CR stack where CR stands for comfy roll the name of the custom node select the CR multicontrol net stack now connect the first pre-processor image output to the image one input of the control net stack node since the previous apply control net node only works with one control net model we need to delete that node next search for nodes using CR and look for CR apply multicontrol net add this node to the canvas now we connect control net stack to control net stack move it up where the other node was connect conditioning to positive and conditioning to negative so this node also uses Nega compared to the previous one then connect base positive to positive and base negative to negative we have a switch here that is turned off if you don't need control net but since I want to use it in this case I'll turn it on delete the load control net model node since it's not needed anymore the stack node already handles it I forgot to remove it so we have the first control net with the cany pre-processor uh we can add another one manually but you can also copy and paste the first control net node the second control net uses the same image but is set to depth allowing us to extract more information and details connected to the second image input in the stack from the portrait image processed with canny get a preview which will be the first image in the stack then process the same portrait with depth add an extra preview to see how it looks and connect it to the second image in the stack now let's review the options on the stack node right now the first switch is off so I will turn it on select the control net model from the list specifically the Promax model we downloaded below you have the strength setting you can adjust the strength start or end at different steps uh allowing control net to apply later in the process and not be so strict if I run this workflow Now it only uses the first control net from the list for the generation to include the second image you need to turn on switch two and select the model for it after doing this running the workflow will apply both control net models canny and depth so the result will be more similar to your input image you can adjust the strength to get more freedom and Variety in the results let's load another image the one with the building I will change the prompt to an apocalyptic building both cany and depth will be set to full strength and this is the cool result I got um it has both depth and details I really like the control that control net offers I save the workflows for both the single and multi-stacked control net so you can download them from my Discord they will be labeled with episode 9 in the name I corrected them after reviewing the edits here's how the save workflow looks I Chang the color to make it clearer select a pre-processor from the list and when you run it you'll see a preview of the pre-processor before getting the final results from load checkpoint make sure you select an sdxl model as the control net Union we downloaded only works with sdxl then add both a positive and a negative prompt in the apply control section you can adjust the strength for the sampler ensure you use the recommended settings for your sdxl model you can also adjust the width and height here next load the control net model and upload an image by clicking the choose file to upload button in this case I have a bunny choose the pre-processor the sdxl model and the resolution after that you can run it and enjoy the control um if you choose a different aspect ratio than the image you uploaded for control net you'll get a cropped version you might see it more clearly if I switch to a wider image so how can you avoid cropping for example I have this wide image of an alien I can try to estimate it by setting a larger width than height with values around 1024 pixels the result is close but not perfect you can adjust these values until you get it right there might be a node that does this automatically but for now I use Photoshop I open the image adjust the width in the image size menu and Photoshop automatically gives me the height then I enter these values into comfy V for width and height which should give me the exact ratio the result is this beautiful alien rabbit creature let's load the second workflow the one with stacked control net I've added some notes here you can turn the switches on or off you can delete the note to make the workflow clear you can also upload an image for control net and choose from different pre-processors in this example I have only two pre-processors activated while the third one is turned off if I copy those two nodes I can create another connection from the uploaded image and choose a different pre-processor like Scribble then I connect that to image three but it's very important to turn on the third one I forgot to switch it on so it didn't pick up the third pre-processor I'm too tired now so I'll finish the video soon but you get the idea and if you need more than three though I've never done that you can duplicate the control net stack add another three images and pre-processors I'll just add one to show you then connect each one to image one image two and so on then I will will remove the connection from the first stack and connect the two stacks together I'll make sure it's set to the one I'm using next I'll make the connection to apply multicontrol net this is the result if the result looks too overprocessed you can try reducing the weight for certain pre-processors to make it less strict and this is the result that's all for today leave a like or a comment if you found something useful see you on the Discord channel for more questions have a great day [Music]

Original Description

In Episode 9 of our ComfyUI tutorial series, we explore ControlNet, focusing on the new Union ControlNet for SDXL. This powerful ...
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
Building the Next Generation of Synthetic Media – OpenCV Live Ep. 218
Learn how generative AI is revolutionizing synthetic media creation, including digital humans and video content, and discover tools like Avatar SDK
OpenCV Blog
📰
Why Diffusion Transformers (DiTs) Are Replacing U-Nets in Generative AI
Learn why Diffusion Transformers are replacing U-Nets in Generative AI and how they improve image generation
Medium · AI
📰
Why Diffusion Transformers (DiTs) Are Replacing U-Nets in Generative AI
Learn why Diffusion Transformers are replacing U-Nets in Generative AI and how they improve image generation
Medium · Deep Learning
📰
How I Built FoodVision Big: Teaching a Model to Tell 101 Dishes Apart
Learn how to build a model that can identify 101 different dishes from food photos
Medium · Machine Learning
Up next
How To Use ChatGPT Image Generator - Easy Way to Get Best Results!
AI Andy
Watch →