This Open-Source AI Video Model Just Crushed Sora
Skills:
Multimodal LLMs90%Generative CV80%Prompt Craft60%Advanced Prompting60%Prompt Systems Engineering60%
Key Takeaways
The video discusses the LTX2 open-source AI video generation model, its capabilities, and its potential applications, including text to video, image to video, and video to video generation, with tools such as LTX2, Loras, Nvidia RTX GPUs, and H100.
Full Transcript
A new AI video generation model just dropped that is free, open- source, and can be ran on your own machine with as little as 12 gigabytes of VRAM. It has synced audio, high resolution output, 30 second long clips, and is up to 18 times more efficient than the next best model. Now, if you haven't heard of it yet, it's called LTX2, and it's not only the current best model out there, but it's open- source, meaning you can download it right now and use it completely for free, even if you don't have the best hardware on the market, which is honestly pretty game-changing. Now, in this video, I'll share more details with you, but massive shout out to LTX for open sourcing this and partnering with me to make this video. So, LTX2 is an open-source audio video generation model released with full weights, training code, benchmarks, and multimodal pipelines. Architecturally, it's a diffusion transformer hybrid designed to generate synchronized video and audio with native 4K output and consistent temporal motion. Now, I know that's a lot of words, but those are the general stats in case you're familiar with these models. Now it supports text to video, image to video, video to video and audio condition generation. It also ships with Loras and control adapters for motion structure and camera behavior. For example, you can have some looks at the website here. You can see 50 SPF performance, right? Native 4K audio and video synced, which is a really big deal. So it actually does both the generation and like syncs it to the lips for example or the sound effects. And then if we have a look here, you can control all of the motion, the structure, the camera, all of that kind of stuff. It's very advanced, and I will show you some examples of what it's generated in a minute. Now, LTX2 is also optimized for local inference on Nvidia RTX GPUs. Now, this includes consumer cards, meaning that you can run it on your own cards with as low as 12 GB of VRAM. For example, a 3060, which was released 5 years ago, you can use that to run this model. Now, the recent release includes the full training recipes, evaluation benchmarks, and also the comfy UI workflows, which enables local inference, fine-tuning, and integration into existing production pipelines. The most insane part here is that this is free, it's open- source, and people have already created some insane stuff with it. So, I just want to go through some of the benchmarks, you can understand how revolutionary this actually is, and then we'll go into a bunch of examples of things that people have built, the GitHub repo, all of that stuff. So in terms of benchmarks, they've released a full research paper in case you want to read more about it. But the one thing to highlight here is that this is the most efficient audio video generation model to date, at least that I'm aware of. And you can see that if we compare it using data center GPUs like an H100. It is 18 times faster, approximately 18 times faster than the next best model, something like WAN 2.2 with 14 billion parameters. You can see just how much faster it is. 49 steps versus 2.69 69 steps per minute on that GPU. And just in my experience, using the API that they provide here, I'm able to generate 10-second videos in about 15 20 seconds, which is extremely fast. And again, you can run this on your own hardware. So, of course, if you're running an older GPU, it's going to take a little bit longer, but even on something like a 5090 or 4090, if you had that locally, you can generate video very, very quickly using this model because of how efficient it is. Okay, anyways, enough of that. I know you guys want to see more examples even though I've played a few on screen. Let me show you what some people from the community have created with this because it's kind of ridiculous. So, I've got a collection of Reddit posts here from the community of people who have played with this model locally and generated some really cool clips. Everything you're going to see here was fully generated with LTX2. Now, some of these people combined multiple clips together, like added some music, but all of the audio in terms like the sound effects or the voice is coming from the model. And of course, obviously, all of the video as well. And in this case, this guy was running this, you know, 16 gigabytes, 64 GB of normal RAM on his own computer. So, he's running it locally. He's not using their API. So, let me run this and I'll just show you a few examples. We have maybe like three or four minutes of super cool stuff. Then, I'll show you the GitHub repo and show you some of my examples that I generated. >> Close-up shots are fine. We've seen that before. But what about the wide ones? >> Well, it may have some limits. It's honestly kind of crazy that with just 16 gigs of VRAM, you can already generate like a 15-second HD video. Yeah, it's not perfect, but look at the motion in the lip sync. It's actually really well done. >> And that ending part with like the girls in the mirror selfie is like crazy to me how accurate that looks. Obviously, there's a few small things like with the hands and how close they are together, but generally the fact that this dude was able to just generate this on his own computer is like ridiculous. Let's look through a few others. There's like some anime things here. >> I mean, I'm not a big anime guy, but you could get the idea. You could make that kind of style or they have like something like this. Let's look at this. >> I will never forgive you for this. >> [music] >> I had no idea DDR prices will go up so high. >> Of course, this guy's just joking around within the audio track. >> How am I supposed [music] to run LTX2 now? >> Please, I will sell a kidney. >> So, I'll play the whole thing, but notice like this is a 27 second long clip where it's all consistent, which usually is not possible with a lot of these video generation models. Usually, they do like 5 seconds, 10 seconds max. I don't know how long it took him to generate this, but 27 seconds is crazy that he was able to get that done uh with the model. Let's look at this one. 53 seconds, right? Same thing. Let's run this. >> Lawrence, >> give me a G note with a fifth above it and the the middle one. No middle one. I changed my mind. Now go an octave below. Now give me some rhythm. And keep that same rhythm. Go. Okay, Katie, remember that note I taught you? The G play, but also keep it rocking good. Okay, give me like a like [music] a jump. Good. Okay, now they're able to do this. Really light. Oh, that's it. Okay, keep going with that. Zack, you remember this thing I [music] taught you a minute ago? Go like Yes. Yes. Wa. All right, [screaming] let's go. >> Like, how crazy is that that you can generate that with AI? So, he did four times 20 second clips. So, he did four clips and stitch them together. But notice he has the consistency between all of them like using the same characters, the same visual style. Um, using a kind of a comfy UI uh what is it? Workflow. This is obviously pretty advanced how he's able to get this, but insane. And then last one, something like this. I actually haven't watched this on. I'm moving. Cover me. >> Say it again. Confirm the address. If I'm wrong, I disappear. If I'm right, if it ends tonight. >> And to be honest, that one's not that good. I think this one is the most impressive one. I'm going to leave these in the description if you guys want to check them out. But absolutely insane how far this these models have come. Like I cannot believe this wasn't produced in a studio and this is like literally just a model doing this. So, I think it's safe to say that this is a pretty ridiculous model, but I know a lot of you are asking, how do we run this ourselves? Now, they have everything completely open source. So, I have the GitHub repo, which I'll link in the description, and it explains here how to set it up. It's pretty much as simple as you need to download the entire model, which I believe is about 300 GB from something like HuggingFace, and then you have to set it up inside of something like Comfy UI. Now, I'm not going to go through the full install here because I'm not an expert at running these models locally on my own computer, but it's all explained in the GitHub repo here. They even have an integration directly for Comfy UI which if you've ever run these type of models locally you're probably familiar with and there is a full tutorial that shows you how to set this up. Now when you're using these models locally typically what you're going to be doing is creating this kind of comfy UI setup where essentially you have a pipeline where you have different inputs and this is how you're able to keep some of like the consistency for example between the images depending on what you're trying to create. So there's a full tutorial here that I'll link in the description. You could also just look it up, you know, how do I run this model locally, but it kind of explains how to set it up inside of Comfy and how to get these kind of pipelines, how to set all of the parameters, so you can fine-tune this and get it to do exactly what you want. I even heard from someone who's on the LTX team that they actually have their own Premiere Pro plugin where they're able to extend a still image or video clip by simply running the plugin directly inside of Premiere, which again you could do yourself. that's not publicly available right now, but kind of insane that you could do that because the model is just out there for anyone to use. So, this is the GitHub. Again, if you want to run it, it is a little bit complex. Does require some kind of understanding of this Comfy UI, but if you want to try it, I'll leave those links in the description and you can mess around with it. Now, of course, another way that you can use this is directly from their API, which I'm going to show you some examples of in a second. Now, the whole point of them making this open source is so you don't need to use the API and obviously you don't need to pay for it or be limited with credits. But if you don't want to go through the whole setup process, then you can of course mess around with it there. Now, at this point, I believe most of you probably think this is cool, but you're probably wondering like why does this really matter? What's so important about it being open- source? And really, the reason has to do less with like the raw quality and more to do with the control of these models. So, most video AI today lives behind a closed API, right? You send a prompt, you get a result, and that's pretty much all that you can do. You can't inspect the model, you can't fine-tune it, and you can't integrate it deeply into an existing production pipeline. Now, with LTX2, the entire system's open, right? The weights are open, the training code is open, and it runs on premise on your own hardware. Now, this changes how this can actually be used. Like, if you're running a studio, for example, or a VFX team, or even your solo creator with sensitive IP, nothing needs to leave your machine. There's no bandwidth limits. no usage caps and there's no blackbox behavior because it runs locally. It can also be embedded directly into real tools like I was talking about having an extension inside of Premier Pro or Da Vinci or something, right? You could have it in a workflow in like Blender or the nodebased systems like Comfy UI. And that's the shift here, right? This isn't video AI as a demo or a toy. It's real video AI as infrastructure, something you can adapt, extend, and actually ship, which is much different than the video AI models that exist today. So, it's pretty cool. You know, massive shout out to them for open sourcing this. It's been, you know, really well received by the community clearly. Now, what I'm going to do is just show you a few examples of some random things that I created that are nowhere near as impressive as these Reddit ones, but that show you like it actually works even with a really simple prompt and someone who's not an expert at using these models. So, the tool I'm inside of right now is called LTX Studio. I believe you do need to pay for this or at least you have like some limited credits. You can see my credits up here. And here you can use all kinds of different models, including obviously LTX2 Pro, LTX2 Fast, some of the other models. They're legacy ones. And there's all kinds of different controls that allow you to generate videos. So, this is the easiest way to play around with it, uh, which is what I've done. So, let me show you a few video clips. These are like photos and stuff that I've taken cuz I wanted to test it on mine. So, this started from like a photo of me on the beach, and you guys will be able to hear the audio even though I'm kind of talking over it. So, you know, it's not perfect. Like, my hat just kind of disappears. Same with my backpack. The interesting thing is like the splash sound effect, right? And the stuff that you get there is pretty cool. So, that's one. Um, let's do some more here. I have like some photos I generated of like Dora. So, I just put a picked a Dora character and said like put her in a desert. Then I did like a video of a watch. This is a photo. And then I just said, "Hey, make it like spinning around the watch here as you can see." Which looks pretty cool. Here's one of like Doris screaming in the desert. >> Help. Is there any water around here? When can I get water again? And then I made a bunch of other ones. Here's one that I thought was pretty funny. >> They can't even make the playoffs. They need a new coach. That's what they need. >> IT'S A DEFENSE. >> It's the forwards. THEY CAN'T SCORE. >> Kind of make memes. So, you can see this is 1440p 10-second video, right? Uh, and then I made some videos of me. So, I just did some like different versions. This is kind of funny. Subscribe. >> Yeah. So, anyways, you get a bunch of stuff. Like I was just playing with it. that I didn't do anything crazy, but there's all kinds of stuff you can do in the studio here, but obviously the more impactful thing is that you can run this model locally on your own computer. So, anyways guys, that's going to wrap it up. This is super cool. It's crazy to see how far we've gone with the AI video generation. I'm definitely going to try to get this up and running on my own machine after downloading the 300 GB and start maybe even using this for some YouTube clips and YouTube videos because it is just really unique. Let me know what you guys think of it in the comments down below and I will see you in the next video.
Original Description
A new AI video generation model just dropped that is free, open source and can be ran on your own machine with as little as 12GB of VRAM. It has synced audio, high resolution output, 32nd long clips, and is up to 18 times more efficient than the next best model. Now, if you haven't heard of it
yet, it's called LTX-2. And it's not only the current best model out there, but it's open source, meaning you can download it right now and use it completely for free.
Want to make real money with coding? I share high-signal insights on careers, monetization, and leverage in my free newsletter. Join here and get my guide How to Make Money With Coding instantly: https://techwithtim.net/newsletter
🎞 Video Resources 🎞
Check out LTX-2: https://ltx.io/model/ltx-2
ComfyUI Setup Tutorial: https://www.youtube.com/watch?v=d1tjLXsz8Wc
Open-Source GitHub Repo: https://github.com/Lightricks/LTX-2?tab=readme-ov-file
Reddit Examples
https://www.reddit.com/r/StableDiffusion/comments/1qae922/ltx2_i2v_isnt_perfect_but_its_still_awesome_my/
https://www.reddit.com/r/StableDiffusion/comments/1qdl0dd/ltx2_vs_wan_22_the_anime_s[…]m=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
https://www.reddit.com/r/StableDiffusion/comments/1q6m285/ltx_is_actualy_insane_musi[…]m=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
https://www.reddit.com/r/StableDiffusion/comments/1qb2cfz/i_recreated_a_school_of_ro[…]m=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
https://www.reddit.com/r/StableDiffusion/comments/1qflkt7/ltx_2_is_amazing_ltx2_in_c[…]m=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
https://www.reddit.com/r/StableDiffusion/comments/1qflkt7/ltx_2_is_amazing_ltx2_in_c[…]m=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
⏳ Timestamps ⏳
00:00 | LTX-2 Model
00:40 | Stats/Specs
02:22 | Benchmarks
03:25 | Reddit Examples
07:09 | GitHub Repo/Running Locally
09:09 | Why this Matters
10:45 | My Generated Videos (LTX Studio)
Hashtags
#LTX2 #AI
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Tech With Tim · Tech With Tim · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
A* Path Finding Algorithm(Visualization)
Tech With Tim
Python Programming Tutorial #1 - Variables and Data Types
Tech With Tim
Python Programming Tutorial #2 - Basic Operators and Input
Tech With Tim
Python Programming Tutorial #3 - Conditions
Tech With Tim
Python Programming Tutorial #4 - IF/ELIF/ELSE
Tech With Tim
Python Programming Tutorial #5 - Chained Conditionals and Nested Statements
Tech With Tim
Python Programming Tutorial #6 - For Loops
Tech With Tim
Python Programming Tutorial #7 - While Loops
Tech With Tim
Python Programming Tutorial #8 - Lists and Tuples
Tech With Tim
Python Programming Tutorial #9 - Iteration by Item (For Loops Continued...)
Tech With Tim
Python Programming Tutorial #10 - String Methods
Tech With Tim
How to Overclock a NVIDIA GPU
Tech With Tim
Python Programming Tutorial #11 - Slice Operator
Tech With Tim
Python Programming Tutorial #12 - Functions
Tech With Tim
Python Programming Tutorial #13 - How to Read a Text File
Tech With Tim
Python Programming Tutorial #14 - Writing to a Text File
Tech With Tim
Python Programming Tutorial #15 - Using .count() and .find()
Tech With Tim
Python Programming Tutorial #16 - Introduction to Modular Programming
Tech With Tim
Python Programming Tutorial #17 - Optional Parameters
Tech With Tim
Python Programming Tutorial #18 - Try and Except (Python Error Handling)
Tech With Tim
Python Programming Tutorial #19 - Global vs Local Variables
Tech With Tim
Python Programming Tutorial #20 - Classes and Objects
Tech With Tim
Cool VBS Script to Prank Your Friends!
Tech With Tim
How to Overclock an AMD GPU
Tech With Tim
Best GPU'S For Mining Ethereum (2018)
Tech With Tim
Recursion and Memoization Tutorial Python
Tech With Tim
Ethereum Mining Rig - Hardware Guide
Tech With Tim
Pygame Tutorial #1 - Basic Movement and Key Presses
Tech With Tim
How to Install Pygame (Windows 8/10)
Tech With Tim
How to Trade Your Cryptocurrency (Bitcoin, Ethereum etc.) For Cash!
Tech With Tim
How to Mine Ethereum 2018 - WORKING (Super-Easy)
Tech With Tim
Microphone Comparison - $10 Mic vs $150 Mic (Blue Yeti USB)
Tech With Tim
Pygame Tutorial #2 - Jumping and Boundaries
Tech With Tim
Pygame Tutorial #3 - Character Animation & Sprites
Tech With Tim
Pygame Tutorial #4 - Optimization & OOP
Tech With Tim
OBS Studio Tutorial - Best OBS Settings
Tech With Tim
Linear Search Algorithm - Python Example and Code
Tech With Tim
Make Any Mic Sound AMAZING! (WITH OBS)
Tech With Tim
Binary Search Algorithm - Python Example & Code
Tech With Tim
Pygame Tutorial #5 - Projectiles
Tech With Tim
Pygame Game - Mini Golf
Tech With Tim
Pygame Tutorial - Projectile Motion (Part 1)
Tech With Tim
Pygame Tutorial - Projectile Motion (Part 2)
Tech With Tim
Pygame Tutorial #6 - Enemies
Tech With Tim
Pygame Tutorial #7 - Collision and Hit Boxes
Tech With Tim
Pygame Tutorial #8 - Scoring and Health Bars
Tech With Tim
Cloud Mining vs. Hardware Mining - 2018
Tech With Tim
How to Install Pygame on Mac OSX (Fast-Simple)
Tech With Tim
Pygame Tutorial #9 - Sound Effects, Music & More Collision
Tech With Tim
Pygame Tutorial #10 - Finishing Touches & Next Steps
Tech With Tim
How to Fade Your Screen in Pygame [CODE IN DESCRIPTION]
Tech With Tim
How to Create a Button in Pygame [CODE IN DESCRIPTION]
Tech With Tim
Pygame Side-Scroller Tutorial #1 - Scrolling Background/Character Movement
Tech With Tim
Pygame Side-Scroller Tutorial #2 - Random Object Generation
Tech With Tim
Pygame Side-Scroller Tutorial #3 - Collision
Tech With Tim
Pygame Side-Scroller Tutorial #4 - Scoring and End Screen
Tech With Tim
How to Create A Message Box in Python - Tkinter
Tech With Tim
Is Ethereum Mining Still Profitable - Is It Worth It (April 2018)
Tech With Tim
How to Run MAC OSX on a WINDOWS PC (Clover Boot-loader)
Tech With Tim
Programming Problem #1 - Alphabet Soup (Beginner/Novice)
Tech With Tim
More on: Multimodal LLMs
View skill →Related Reads
📰
📰
📰
📰
The Best Free AI Image Generators Better Than ChatGPT and Gemini
Dev.to AI
50+ Sequential Images, One Prompt in Codex
Medium · ChatGPT
How can I batch-generate 3D assets from prompts or images using an API, and which 3D generation APIs support batch generation?
Reddit r/artificial
How AI Head Swap Works: The Technology Behind Realistic AI Image Replacement
Dev.to AI
Chapters (7)
| LTX-2 Model
0:40
| Stats/Specs
2:22
| Benchmarks
3:25
| Reddit Examples
7:09
| GitHub Repo/Running Locally
9:09
| Why this Matters
10:45
| My Generated Videos (LTX Studio)
🎓
Tutor Explanation
DeepCamp AI