ComfyUI Tutorial Series: Ep13 - Exploring Ollama, LLaVA, Gemma Models

pixaroma · Beginner ·🧠 Large Language Models ·19:54 ·1y ago

Key Takeaways

This video tutorial series covers using Ollama with ComfyUI to run and test LLaMA, Gemma, and LLaVA models.

Full Transcript

welcome to episode 13 of our comfy UI tutorial Series today I want to show you how to install and run the Alama tool and how you can test different large language models inside comfy UI with its help this way you can generate prompts from your instructions or even from images first we need a tool to manage all those large language models so go to the ama.com website as you can see it can manage all kinds of models like llama fi uh Mistral and lava uh if you want to learn more about it there's a docs section that links to the GitHub page you'll also find more uh links here for a quick start and additional resources back on the AMA site you have the download button where you can choose your operating system mine is Windows so I'll click download for Windows I'll save it on my desktop for now and let's see how we install it the icon features a cute llama just double click on it and then click install and wait for it to finish once it's installed you can see the Llama icon in the notification area that indicates it is running if you want to close Al rightclick there and choose quit it's important to know this as it can use up some vram so you can close it when you're not using it to start Alama again find it in your apps by clicking on start and searching for AMA now if I click on it it might seem like nothing happened but llama has started as we can see the icon with the llama in the notification area back on the Yama website we have a section for models if I click on it we can find different models here such as llama Gemma mistol and so on lots of options to choose from the top search is a little better than the filter below so if I search for vision for example I can find models that can interpret an image and provide a description or caption for it if I search for Gemma I can see different variations of that model if I go back to the models we can also filter them by newest or uh popular here are some popular models let's say I want to install and test Gemma I can see different versions available if I click on it and scroll down I can read more information about it here you can see different versions but to see all you need to click on uh view more as you can see it has all kinds of sizes usually the larger models are smarter but it also depends on how many parameters they are trained on for example 2B means 2 billion parameters so 7B has 7 billion parameters which allows it to handle more complex tasks but requires more processing power if I select the 7B version I see it a command on the right to install that model if I choose a different model like 2B instruct Q4 and the command changes I can click on this icon to copy that command then I click Start type CMD and select command prompt here I can paste the command I just copied and press enter this will pull all the information and models it needs to run this model has less than 2 gigabytes so it should run quickly but it isn't as smart as some other models however there is also a Gemma version two that seems to be better once the model finishes installing we get a prompt asking us to send a message as suggested you can type a message to get help and give some command suggestions like for slby for exit the forward slel command also works to show all those commands we can communicate with the model right here so if I type high and hit enter it will respond it's nice to interact with the model but the interface is kind of basic there are graphical interfaces like open web UI but today I'll focus on using it uh inside comfy UI I'll use forward slby to exit if you use the command Al list you can see all the models you have installed as you can see I have a few gigabytes here which can take up a lot of space on your hard drive so if you want to remove models you don't use anymore select the name of the model you want to remove like stable prompt for example then use contrl plus C to copy that name type Al followed by RM which stands for remove paste the model name and press enter that model will be removed if I run the list again you can see it's not in the list anymore I'll close this window and now let's go to comfy UI let's go to the manager and then to custom nodes search for AMA and install the comfy UI olama node by this user click install and wait for it to finish then you can click restart and okay we can close this tab now since a new window will open clear the canvas using the clear button then double click on the canvas and search for Al select the AL generate node this one has a text output called response and we can add a show any node from comfy easy use custom nodes here you can add your instructions or text for example if I say hi and run it I get this response this is using the URL that points to AMA running in the background if I copy this link and paste it in the address bar you can see it says Alma is running so we need Al to be running to use the models you can give all kinds of instructions like in chat GPT now it's not always perfect it depends on how smart the model is and how many billion parameters it has has but it's free and you can choose a model that fits your computer's power allowing us to generate prompts with it let's delete the AL generate node and add another one called Al generate Advanced um this one has more options connect the response output to the show any node don't use the context output as that will give a bunch of numbers now here we have the instructions for what we want to do and at the top we can place our prompt uh when we run this it will interpret that prompt using the instructions you can try all kinds of instructions to see what works best for you I also tried generating some instructions with chat GPT I will post these long instructions for generating a prompt that works well most of the time though it depends on the model when I was trying to get the prompt to start with the type of image I mentioned watercolor to see if it would start with that however I need to try it with some larger models for better results or to play with settings let's try this version to see if it works better from here you can take for example the 7B version which is better or look at Gemma version 20 with 7 billion parameters but just to show you how to add another model I'll demonstrate with this 2B version which I think is similar to the Q4 version though I'm not sure anyway copy the command go to start search for CMD and uh open the command window paste the command and wait for it to install then close it with the command bu or just close the window now back in comfy UI it doesn't appear here if I try to refresh I get undefined model so you can either restart comy UI add the node again or right click and choose fix node which will recreate that node with default values there are ways to make it more accurate if you want to experiment for example if I take a screenshot of this node and go to chat GPT and paste that screenshot it only works with a premium subscription by the way so I can ask about what those options mean like top K temperature keep alive and so on I can also ask what settings to change to get a more accurate prompt for stable diffusion here's what was suggested to lower the value of top K and top PE and increase the temperature among other adjustments you can play around with these settings to see if you can find better options let me add the Gemma 7 billion parameter model as well you know the drill copy the command and run that in the command window back in compy UI it doesn't appear since it was open so I'll just rightclick and recreate the node uh now I have that 7B version there I will paste the instructions and then change the prompt and uh this one works much better I'll probably find even better versions after I finish the video so make sure to check Discord to see what people have found that works best I had a problem with the Discord link so the invite link from the header where you usually subscribe to the channel should work fine um let's remove this uh node and add one called AMA vision uh we can connect the description output to the show any node and then on the left we can add a load image node we had something similar in episode 11 uh now when I run this workflow it says it is not able to access external information or visual content so this model is good for text but cannot read images let's find a model that can do that let's search the Alama website for the word Vision here um you can find a lot of vision models that can interpret an image but many of those are variations of the lava model uh if you want a smaller model you can try something like lava 53 for example um you can read the description there are more versions um like the speni version so let's test it to see how it works run that command in the command window then back to comy UI restart comfy UI or recreate that node now we can select that model from from uh the list when we run it we get a uh a description of the image I can ask questions about that image like what color the cube is let's try with a different image like this portrait I can ask about her hair so I can use that simple describe the image instruction and I will get a prompt the better the model the better the prompt will be let me try a bigger model like the lava one I will try the 7 billion parameter version which has has almost 5 gbt copy the command and install it using the command window if I run the AMA list command you can see I've already tested a few and need to try even more when I get more time if you find some really good models let me know in the comments or on Discord back to comfy UI let's recreate the node so we have access to the model we just downloaded I will select it from the list and test it the result is a really detailed prompt now this Alama Vision node doesn't seem to have a seed function so if I try to run it multiple times The Prompt doesn't change unless I change the photo or the instruction what you could do is add a space or a comma in the description so the instruction changes um that way it will let you generate another prompt as you can see it generated this long prompt if you want another prompt change the instruction Again by deleting the space you added or adding another space or anything else that changes it um if you know a better way let me know um for more prompt models not Vision but those you use for text you can also try the FI version like fi3 mini which are quite small um depending on your video card you can also try to integrate it with a workflow for sdxl models it worked quite well for me but on the flux model it struggled so let's add the Alama generate node let's use a normal model not a vision one like Gemma 7 in this case and then I will add a show any node just to see what prompts it generates now if we try to connect the positive prompt we cannot so we need to add that input right right click select convert widget to input and then convert text to input now we can connect the AL generate to that node let's add some custom instructions and then a prompt that will use those instructions the result is this cute bunny um I will paste those long instructions for the prompt generator and let's try again now it's more detailed we can add more information to the prompt to see if it can improve it uh now we have more information there about the style as well the result is quite nice let's see if it knows how to do a watercolor um by the way flux models don't know as many art styles as the sdxl models so I hope they will fix the model in the future let's try a 3D render also and it seems to know how to do it so it can be quite useful um let's try another one on a black background and seems to work okay Gemma 7B might be too big for some video cards so try smaller models if your video card cannot handle both or just generate prompts first paste them in a text file then shut down the Alama app from notifications and you can use your workflow normally since running both Alama and comy UI at the same time can take more resources here I converted the system field the one we use for instructions to a text input so I can add a positive node and con it to that input then it will work the same but you have a bigger window where you can work more easily on your instructions if you change those more often I will post my long instructions here you can see I tried to guide it so it knows uh how to prompt better for stable diffusion and it worked okay most of the time the result is this bunny now let's delete all these nodes so we can test the vision models as well well add the AL Vision node then the load image node where we load our image and then a show any node so we can see what prompts it generates now let's connect it to the workflow for the model we need to change it to one of those Vision models so I will go with the lava model now it describes that image and uses it as a positive prompt generating an image based on that prompt in this case it worked quite well let's try it with a portrait I got a nice result with the the portrait but I wanted a darker background so let's change the instruction by adding a space to generate a a new prompt I tried a few times but had problems identifying that background color um probably because I saw some lights there in the background I thought why don't I add a black background in the instructions so I added make the back ground black and um in this case it worked how I wanted the images are quite similar let's try a vector illustration of the cat um this one worked okay now but it didn't work all the time um sometimes I had to specify in the instructions that I wanted it in Vector art style for the anime image I got a photo image instead even though anime was mentioned in the prompt for that flux is better at understanding prompts but it recognizes fewer art styles compared to sdxl so I added more anime related words in the instructions to guide it to what I wanted in the end adding anime 2D and chibi seem to help achieve the results I wanted this is how the final workflow for sdxl looks um you have some ideas for instructions and models here what I want to show you is um how to easily move part of this workflow to another existing workflow so for example if I want to use these four nodes I select them by holding control then rightclick on the canvas and choose save selected as template you can give it a name like Alama Vision or something you'll remember now we can open another workflow like this chel one for example um for Dev it worked really slowly For Me Maybe it's too much for both Al and flux Dev anyway now if we right click and go to node templates we can select those save nodes and add all of them to this workflow if I hold shifts I can move them all together where I want them now remember we can't connect directly we need to convert text to input and now it lets us connect that's all now we have the Alama Vision in this workflow as you can see it runs uh for flux Dev it got stuck at 5% in K sampler many times but this one worked okay okay this time it missed some detail so I will add them in the instructions to help it get a better description and it worked acceptably now for those big workflows that take a lot of time I prefer to First generate the prompts or text using a small workflow then I copy that text to a text file and load the flux workflow to try those prompts with these workflows you can give custom instructions and text so you can either have some fun like asking it to act like a pirate or you can request a story with the same instruction whether it's about a pirate or any character you want I also use it for generating ideas uh before I create any prompt I ask for a list of ideas to see what I could try if I like an idea I can load a flux workflow and shut down theama app alternatively I can add a promp generator workflow to get a better prompt uh allowing for more creativity um see have created a workflow specifically for getting prompts from images where you can try larger models to see what works uh and get a more accurate image descriptions um then you can close that and the Alama app uh use flux Dev to generate an image a heartfelt thank you to everyone who joined as YouTube members a special shout out to the Legends for all your support as well as to the VIP and other members on Discord who helped this channel Thrive if you found this video helpful please leave a like or a comment thank you again and have a fantastic day [Music]

Original Description

In this episode, we'll show you how to use the Ollama tool with ComfyUI to run and test models like LLaMA, Gemma, and LLaVA.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

This tutorial teaches how to use Ollama with ComfyUI to run and test popular LLM models like LLaMA, Gemma, and LLaVA. It covers the basics of using the Ollama tool and ComfyUI for AI model testing. By following this tutorial, viewers can learn how to work with different LLM models and integrate them with ComfyUI.

Key Takeaways
  1. Install ComfyUI
  2. Set up Ollama tool
  3. Run LLaMA model
  4. Test Gemma model
  5. Integrate LLaVA model with ComfyUI
  6. Configure model settings
  7. Run and test models
💡 The Ollama tool can be used with ComfyUI to easily run and test different LLM models, making it a useful tool for AI model development and testing.

Related Reads

Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →