10x Faster Than Standard LLM!? DiffusionLM Explained

bycloud · Beginner ·📄 Research Papers Explained ·11mo ago

Key Takeaways

Explains DiffusionLM, a technique for improving LLM performance, achieving 10x faster results than standard LLMs

Full Transcript

AI based applications currently have a major problem that is an insanely long wait time. Not only does ALM now think which can take a while to get an answer, but when incorporated with other AI agentic applications, the weight is even more painful. And the only few solutions we currently have right now is either use a dumper model or just have better hardware to generate faster, which most of us can't control. But there is an architecture that is now on the rise which shows promises that might just solve the speed problem maybe once and for all and that is the diffusion language model which you may have remembered from my video around half a year ago. And oh boy has the evolution been glorious. So back in February a company called Inception Labs published the first ever commercial ready the Fusion LM that you can try out. And for its Mercury coder mini model, it can already generate up to 1,19 tokens per second, which is five times faster than Gemini 2.0 flashlight and quen 2.5 coder 7B, which are some of the top performance small coding models at that time. A staggering speed that auto reggressive models just cannot comprehend. Then surprisingly, the next company that hopped onto diffusion LM is actually Google DeepMind where they unexpectedly announced Gemini diffusion during Google's annual conference this year. You can try it out right now by signing it up to their weight list and I got off pretty fast to be honest and just look at the generation speed. That's even faster. I did not fast forward anything. This is literally how it looks like. So in this video we will take a look at how realistic can diffusion LM measure up against auto reggressive LLM, how it could overcome the reverse curse and take a peek at multimodal diffusion LM which might actually be a very great idea. Before we dive into it, let me just quickly share with you about Worp. An agentic development environment like Claude Code, but with a built-in code editor, semantic search, and full agent control across multiple repos. So, think of this as an AI powered coding platform like VS Code or Cursor, but with no clutter and agents that can handle multi-step workflows. Instead of writing boring commands, you can simply describe your goal in English and Warps agent would take care of the rest. Here, I have my repository for a new feature that I'm building for my website findmy papers.ai with a handful of files and only a tiny bit of documentation. But instead of digging around manually to refresh my knowledge, if I type analyze this repo, Warps Agent will parse the photo structure, inspect key files, and present a concise overview. So for my newest function called scout, which is a custom semantic archive alerts for topic of your choice, there is this feature that I need to build where I need to screenshot the first page of the PDF and present it as an image thumbnail within the scout page. As I describe what I need pretty simply, warp was able to find my existing placeholders for these thumbnail functions that are in the database, my utils, and my files. collecting all this information together and basically implement it for me. For every new line of code it generates or changes, warp will show diff and ask for permissions to write and execute. You can also make it auto execute or even turn on advance planning using 03. But since auto execution is something risky, they have guard rails to catch sensitive commands. And you can also define the guar rail aka the deny list yourself. And while one of the agents is working, you can start another in parallel to work on another task. Here I wanted to write some test files for my existing function that extracts the affiliation of a paper. And you can also see it's proposing me if I want to use 03 to generate a plan. In the meantime, you can also add workflows where it can hold complicated commands so you don't have to manually type it every single time. Here I set it to execute multiple existing test scripts that I have and check if everything is still working as expected in one click. And I can just simply drag this workflow from my personal workspace into the team space on my warp drive so my team can also access it. And beside workflows, you can also add custom rules which is kind of like a system prompt. So other than telling warp to add mu at the end of every sentence, you can also give more context like the virtual environment you use, what OS you are on, etc. And to double check the implementation, I have this pipeline I quickly drew before. Worp can understand the image and check if they implement it correctly. And the best thing is when you want to fully embrace vibe coding and have it test something until it works, you can just turn on the auto execution on the bottom right, then sit back and relax, grab a coffee, and come back to an agent completing a feature for you. So, if you want to try out using Warp 2.0 for just $1 in your first month, enter the code by cloud at checkout to get basically 95% off. And don't forget to use the link down in the description. Thank you Warp for sponsoring this video. Anyways, unlike auto reggressive models we currently all use, which generate tokens from left to right, diffusion language models start from a fully masked sequence and denoise it over multiple iterations similar to image diffusion, but operating on discrete tokens instead of continuous pixel values. At each step, the model predicts a subset of the mask positions in parallel guided by a masking schedule or confidence threshold and replaces them until the sequence is completely unmasked. So this non-auto regggressive process that has a birectional or even the entire global context makes it great at doing structured or logical tasks. It might sound like diffusion LM is a diffusion model, but technically it's actually a method. The architecture it relies on is still transformer. Under the hood, the entire mass context is fed through a transformer that implements the den noising process. So you simply iterate it until the sentence fully emerges, which is referred to a course to fine generation process. So why does the existing chat interface from Gemini diffusion and inception labs still seem like it is generating from left to right? Well, in most cases, it's just an effect that you can toggle. You really can't trust anything you see online nowadays, can you? But there are semi-auto reggressive methods that combine the left to right properties with diffusion LM. But I don't want to get too technical in this video. So if you guys want a more in-depth video on how diffusion LM works, let me know down in comments or you can ask my web app find my papers.ai, my chatbot for AI papers. So with commercial diffusion LMS finally in the wild, we can at last compare them against auto reggressive heavy weights and see how they actually measure up. For inception labs, they have officially published two models back in February. One is called Mercury Coder Small and the other is called Mercury Coder Mini. Although they only officially serve Mercury Small through their API or their web chat. It is a 32K context window model with a price sitting at 25 cents per million tokens input and $1 per million tokens output. According to Open Router and Artificial Analysis, Mercury has around 800 tokens per second generation speed on an H100 GPU. In contrast, the fastest autogressive service clocks in at around 331 tokens per second and engineering fees like Cerebra's custom hardware for Llama for Maverick hit around 2522 tokens per second, but only on their ultra expensive silicon. Looking at their official benchmarks, Mercury Coder is primarily being compared to older lightweight code focused models like Gemini 2.0 flashlight and Quinn 2.5 coder 7B in a way hinting that its parameter count might also be around there. Across code generation tasks, Mercury matches or slightly exceeds these baselines. It especially stands out in the fill-in-the-middle benchmark where it outperforms competitors by 22 to 30% in accuracy. Pretty impressive even if it was primarily trained for it. This means that Mercury Coder is pretty much made for coding. But while trying to validate the results, I couldn't find any other thirdparty benchmarks on it other than the co-pilot arena. What the co-pilot arena measures is the quality of auto completion. speed here is actually not a big player because the AB test for results would be displayed at the same time even if a model completes faster than the other. Interestingly, Mercury actually ranks sixth on this benchmark but with a few catches. This arena only has non-reasoning models and most of them are like a year old. So, it being the sixth isn't really representative in the larger picture. The only good bar here to measure against might just be 3.5 SA new and Quinn 2.5 coder. So looking from these two models, it seems like the progress of diffusion LM is around 8 to 10 months behind private models and only around 3 months behind open source models in the coding domain and also for a smaller model size. But Google's latest Gemini diffusion beta is already doubling Mercury on raw speed, peaking at nearly 1,500 tokens per second with identical 32K token context window. Their official benchmarks only compare against Gemini 2.0 zero flashlight. But across those coding benchmarks, Gemini and Mercury both performed similarly. So the real edge for Mercury right now is that it's the only one that has an API available. But what gave some extra nice insights is that Gemini diffusion was also benchmarked beyond just code. On their official benchmarks, Gemini diffusion falls a bit behind Gemini 2.0 zero flashlight in language and reasoning tasks with GPQA accuracy being 16% lower, multilingual QA being 10% lower, reasoning being 6% lower with only a tiny win on AM math that has 3% higher performance. Hopefully, this gap simply reflects its exclusive focus on coding data rather than revealing a fundamental flaw in scaling diffusion LMS. But a model that's capable of answering pretty much at an instant, I am so incredibly bullish on that. That being said, matching outer regressive models on reasoning, especially challenging the current test time compute paradigm, won't be so easy as you can just simply throw more yapping time at them due to its fixed context window. On top of that, curating synthetic reasoning data might still need to be bootstrapped from auto reggressive models to close the quality gap. Also, with things like higher confidence thresholds, spamming more dinoising iterations, larger model sizes, and some hyperparameter sweeps still have spaces left to be explored. With these grounds to cover, diffusion LM might still need some time to overtake traditional auto reggressive dryins, if that's even possible, of course. Not to mention, open source diffusion LMS actually have quite a lot of optimization problems with the KV cache. So, it's not like the one you downloaded now on hooking phase would generate up to 1,000 tokens per second right away. But there is a new research that came out recently which showed that diffusion LM tends to learn more if data is scarce and compute is not. as the same data set can provide 25 times more learning signals to diffusion LM than traditional auto reggressive models. Additionally, what diffusion LM also has a pretty good chance of winning at in terms of performance might be the multimodal department as diffusion LM might be a great backbone for an unified model as it brings a natural advantage when you throw vision into the mix. Two standout examples come from these two research papers called Lavida and Mada which are both published recently showing why multimodal diffusion LM might be a better choice than multimodal auto reggressive models. For lvida which is short for large vision language diffusion model was masking. This paper basically introduced the first family of VLMs based on diffusion models. So instead of processing image features and text left to right, Levida prepends a fixed vision prefix that are just basically embeddings from a frozen image encoder to the entire mask sequence. So every token including image tokens are being attended birectionally to all other visual and textual context at each iteration. You can then do things easily that requires infilling or early coordination. For example, generating a poem where each line starts with a specific syllable or extracting structured information from an image into a predefined JSON format with no extra prompt engineering. So for format intensive things like coding, diffusion LM may just be the way to go, especially with image understanding support. But on the topic of integrating image understanding, if you remember my native multimodel LM video, wouldn't doing it natively, aka early Fusion, still be better than using a frozen image encoder? Coincidentally, Mada, which is also released in the same week as Lvida, proposed exactly this. If you remember that early fusion video, you might remember Chameleon. And Ma is just like Chameleon, but instead of being an auto reggressive method, is now generated with a discrete mask diffusion backbone. And when you have an unified architecture, this diffusion transformer can literally do anything from text understanding, vision reasoning, and even text to image generation is possible. Their open source model AB has language capabilities better than Llama 3AB and near the performance of Quinn27B. It has better image quality than Dolly 2 and SDXL and also outperforming outer regressive based unified models like Chameleon and Deepseek Janis while having comparable results with standalone vision encoders that can only do vision understanding. How it generally works is that images are quantized into discrete tokens then mixed with words in the same mask prediction scheme. During fine-tuning, it is trained with chain of thought format, teaching the model to plan first and then produce results. So this model with no separate decoders, no left to right yapping, is able to hold its own ground against existing specialized autogressive models. And theoretically, the idea of native multimodal diffusion LM just makes so much sense. One is that you basically have global context everywhere. So vision and language tokens will see each other at every step with no directional blind spots. Second, the multimodal outputs can be amasked in batches through the dnoising process. So long image text sequences would finish faster than auto reggressive pipelines. Third, the model would have a more consistent structure control. So infilling a table and formatting a JSON just flows naturally from the mass predicting course to fine generating paradigm. So maybe in the next video I'll talk about the technical side of diffusion LM like how it exactly works the challenges it has with efficiency and how they can be unified while balancing the loss signal from multi modality. So subscribe to stay tuned and if you like today's collection of papers definitely check out my newsletter where I cover the latest and the juiciest papers weekly on there. I have already covered a few of these papers. So if you do not want to miss out on the cutting edge research development definitely go check it out. And thank you guys for watching. A big shout out to Andrew Chellius, Chris Leoo, Degan Gan, News Research, Kanan, Robert Zaviasa, Leis Muk, Ben Shainer, Marcelo, Ferraria, Zane, Sheep, Poof, and Enu, DX Research Group, and many others that support me through Patreon or YouTube. Follow me on Twitter if you haven't and I'll see youall in the next

Original Description

Try out Warp 2.0 now, the current rank #1 AI on Terminal Bench, outperforming Claude Code: https://go.warp.dev/bycloud You can also use code "BYCLOUD" to get Warp Pro for 1 month free. (limited for 1,000 redemptions) My Newsletter https://mail.bycloud.ai/ my project: find, discover & explain AI research semantically https://findmypapers.ai/ My Patreon https://www.patreon.com/c/bycloud Video Sauces: Inception Labs [Website] https://inceptionlabs.ai/ Gemini Diffusion [Blog] https://deepmind.google/models/gemini-diffusion Large Language Diffusion Models [Paper] https://arxiv.org/abs/2502.09992 MMaDA [Paper] https://arxiv.org/abs/2505.15809 LaViDa [Paper] https://arxiv.org/abs/2505.16839 Diffusion Beats Autoregressive in Data-Constrained Settings [Paper] https://www.arxiv.org/abs/2507.15857 Try out my new fav place to learn how to code https://scrimba.com/?via=bycloudAI This video is supported by the kind Patrons & YouTube Members: 🙏Nous Research, Chris LeDoux, Ben Shaener, DX Research Group, Poof N' Inu, Andrew Lescelius, Deagan, Robert Zawiasa, Ryszard Warzocha, Tobe2d, Louis Muk, Akkusativ, Kevin Tai, Mark Buckler, NO U, Tony Jimenez, Ângelo Fonseca, jiye, Anushka, Asad Dhamani, Binnie Yiu, Calvin Yan, Clayton Ford, Diego Silva, Etrotta, Gonzalo Fidalgo, Handenon, Hector, Jake Disco very, Michael Brenner, Nilly K, OlegWock, Daddy Wen, Shuhong Chen, Sid_Cipher, Stefan Lorenz, Sup, tantan assawade, Thipok Tham, Thomas Di Martino, Thomas Lin, Richárd Nagyfi, Paperboy, mika, Leo, Berhane-Meskel, Kadhai Pesalam, mayssam, Bill Mangrum, nyaa, Toru Mon [Discord] https://discord.gg/NhJZGtH [Twitter] https://twitter.com/bycloudai [Patreon] https://www.patreon.com/bycloud [Business Inquiries] bycloud@smoothmedia.co [Profile & Banner Art] https://twitter.com/pygm7 [Video Editor] @Booga04 [Ko-fi] https://ko-fi.com/bycloudai
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from bycloud · bycloud · 0 of 60

← Previous Next →
1 Can Deepfake work on Anime?
Can Deepfake work on Anime?
bycloud
2 AI that Can Copy Voices
AI that Can Copy Voices
bycloud
3 Live Action Is Terrible So AI Turned It Back Into Anime
Live Action Is Terrible So AI Turned It Back Into Anime
bycloud
4 2 AIs Enhance Anime to 4K 240FPS, but is it good?
2 AIs Enhance Anime to 4K 240FPS, but is it good?
bycloud
5 IRL to Anime With Cartoonization AI
IRL to Anime With Cartoonization AI
bycloud
6 How Does AI Generated Songs Sound Like? [OpenAI Jukebox]
How Does AI Generated Songs Sound Like? [OpenAI Jukebox]
bycloud
7 AI Makes Any Images Cinematic [3D Photo Inpainting]
AI Makes Any Images Cinematic [3D Photo Inpainting]
bycloud
8 AI Generates Anime Faces, And It's Getting Even Better [StyleGAN2]
AI Generates Anime Faces, And It's Getting Even Better [StyleGAN2]
bycloud
9 Tech Behind The Meme: Dame Da Ne AI - Single Image Deepfake
Tech Behind The Meme: Dame Da Ne AI - Single Image Deepfake
bycloud
10 AI Generates New Light Source for Images [PaintingLight]
AI Generates New Light Source for Images [PaintingLight]
bycloud
11 Depixelizing Doom Guy? Mona Lisa in Real Life? The "Upscaling" AI: PULSE
Depixelizing Doom Guy? Mona Lisa in Real Life? The "Upscaling" AI: PULSE
bycloud
12 Image Completion AI - Predict Pixels Just Like Text Predictions [Image-GPT]
Image Completion AI - Predict Pixels Just Like Text Predictions [Image-GPT]
bycloud
13 AI Generates 3D Human Model from 2D Image [PIFuHD - FacebookAI]
AI Generates 3D Human Model from 2D Image [PIFuHD - FacebookAI]
bycloud
14 AI Assisted Masking - Save Your Precious Time Right Now [AE Rotobrush 2]
AI Assisted Masking - Save Your Precious Time Right Now [AE Rotobrush 2]
bycloud
15 This AI Reconstruct Real Life Objects From Just Images [NeRF]
This AI Reconstruct Real Life Objects From Just Images [NeRF]
bycloud
16 Image Restoration AI - Upscale and Restore Faces with DFDNet
Image Restoration AI - Upscale and Restore Faces with DFDNet
bycloud
17 Best Image Colorization AI 2020
Best Image Colorization AI 2020
bycloud
18 Image Decomposition AI - Edit Highlights and Textures Easily [Appearance Eraser]
Image Decomposition AI - Edit Highlights and Textures Easily [Appearance Eraser]
bycloud
19 Deepfake With Audio Only [Wav2Lip]
Deepfake With Audio Only [Wav2Lip]
bycloud
20 Copy IRL, Paste on your PC [AR Cut & Paste]
Copy IRL, Paste on your PC [AR Cut & Paste]
bycloud
21 This AI Transform Faces into Hyper-Realistic Cartoon Characters [Toonify]
This AI Transform Faces into Hyper-Realistic Cartoon Characters [Toonify]
bycloud
22 This AI Restores Old Photos with Damages Automatically!
This AI Restores Old Photos with Damages Automatically!
bycloud
23 Anime Filter with AI - Snapchat vs. TikTok
Anime Filter with AI - Snapchat vs. TikTok
bycloud
24 AI Reduces Bandwidth Problems for Video Calls [NVIDIA Maxine]
AI Reduces Bandwidth Problems for Video Calls [NVIDIA Maxine]
bycloud
25 AI Motion Capture - Track Your Hands & Body WITHOUT Bodysuit [FrankMocap]
AI Motion Capture - Track Your Hands & Body WITHOUT Bodysuit [FrankMocap]
bycloud
26 AI Converts Cartoon Characters To Real Life [Pixel2Style2Pixel]
AI Converts Cartoon Characters To Real Life [Pixel2Style2Pixel]
bycloud
27 AI Sky Replacement with SkyAR
AI Sky Replacement with SkyAR
bycloud
28 Better Than DAIN? NEW BEST Tool for Boosting Video's FPS with AI [RIFE/Flowframes]
Better Than DAIN? NEW BEST Tool for Boosting Video's FPS with AI [RIFE/Flowframes]
bycloud
29 AI That Paints Anything Stroke By Stroke
AI That Paints Anything Stroke By Stroke
bycloud
30 What Happens When AI Robots Design Themselves
What Happens When AI Robots Design Themselves
bycloud
31 Deepfake Movements with 1 image ONLY [Liquid Warping GAN]
Deepfake Movements with 1 image ONLY [Liquid Warping GAN]
bycloud
32 ANYTHING can be a "Green Screen" Now [Real-Time High-Resolution Background Matting]
ANYTHING can be a "Green Screen" Now [Real-Time High-Resolution Background Matting]
bycloud
33 AI Transform any Image into Sketch or Line Art [ArtLine]
AI Transform any Image into Sketch or Line Art [ArtLine]
bycloud
34 AI That Could Soon Replace Vector Artists [DALL-E]
AI That Could Soon Replace Vector Artists [DALL-E]
bycloud
35 Photoshop Detector AI Is Useless
Photoshop Detector AI Is Useless
bycloud
36 The Future Of Online Shopping
The Future Of Online Shopping
bycloud
37 How The Future of Image Search Would Look Like
How The Future of Image Search Would Look Like
bycloud
38 Everyone Can Make 3D Animations Easily Now! [Monster Mash]
Everyone Can Make 3D Animations Easily Now! [Monster Mash]
bycloud
39 3D Video Stabilization with AI [NSFF]
3D Video Stabilization with AI [NSFF]
bycloud
40 OpenAI’s Sarcastic Chat Bot [GPT-3 API Beta]
OpenAI’s Sarcastic Chat Bot [GPT-3 API Beta]
bycloud
41 You Describe & AI Photoshops Faces For You [StyleCLIP]
You Describe & AI Photoshops Faces For You [StyleCLIP]
bycloud
42 You Only Need Audio To Deepfake Now! Might look slightly cursed tho [PCAVS]
You Only Need Audio To Deepfake Now! Might look slightly cursed tho [PCAVS]
bycloud
43 This AI Transfers Anime Back Into Sketch [Anime2Sketch]
This AI Transfers Anime Back Into Sketch [Anime2Sketch]
bycloud
44 AI Learns To Play CS:GO By Watching Humans Play!
AI Learns To Play CS:GO By Watching Humans Play!
bycloud
45 How AI Fixes The Horrendous CR7 Statue
How AI Fixes The Horrendous CR7 Statue
bycloud
46 Best Vocal Isolation & Instrumental Extraction 2021 [lalal.ai vs Spleeter]
Best Vocal Isolation & Instrumental Extraction 2021 [lalal.ai vs Spleeter]
bycloud
47 Face Enhance AI Restores Extremely Blurry Faces [GPEN]
Face Enhance AI Restores Extremely Blurry Faces [GPEN]
bycloud
48 AI That Only Needs 1 Image To Deepfake [SimSwap]
AI That Only Needs 1 Image To Deepfake [SimSwap]
bycloud
49 The Amazing AI Behind the TikTok JoJo Pose Challenge [BoostMonocularDepth + 3DP]
The Amazing AI Behind the TikTok JoJo Pose Challenge [BoostMonocularDepth + 3DP]
bycloud
50 StyleGAN3!? - What AI Actually Sees When Generating Faces [Alias-Free GAN]
StyleGAN3!? - What AI Actually Sees When Generating Faces [Alias-Free GAN]
bycloud
51 AI generated art goes brrrrr [VQGAN+CLIP]
AI generated art goes brrrrr [VQGAN+CLIP]
bycloud
52 AI That Doodles Any Given Description
AI That Doodles Any Given Description
bycloud
53 Best AI Motion Capture 2021 - OpenPose vs DeepMotion
Best AI Motion Capture 2021 - OpenPose vs DeepMotion
bycloud
54 Anime Image Enhance AI Has Gone To The Next Level [Real-ESRGAN]
Anime Image Enhance AI Has Gone To The Next Level [Real-ESRGAN]
bycloud
55 This Video's Voice Is Entirely Made From Audio Deepfake
This Video's Voice Is Entirely Made From Audio Deepfake
bycloud
56 I Can’t Sing So I Cloned My Voice w/ AI To Cover Goodbye Sengen (English Cover)
I Can’t Sing So I Cloned My Voice w/ AI To Cover Goodbye Sengen (English Cover)
bycloud
57 Best Background Removal - AIs Removes BG Without Green Screen And It's Amazing. [RVM]
Best Background Removal - AIs Removes BG Without Green Screen And It's Amazing. [RVM]
bycloud
58 How I Deepfaked VTuber Gawr Gura with AI
How I Deepfaked VTuber Gawr Gura with AI
bycloud
59 AI Magic Removal - Removes ANYTHING & Inpaints For You [LaMa]
AI Magic Removal - Removes ANYTHING & Inpaints For You [LaMa]
bycloud
60 I Did NOT Expect AI Anime Filter To Be This Good [AnimeGANv2]
I Did NOT Expect AI Anime Filter To Be This Good [AnimeGANv2]
bycloud

Related Reads

Up next
Welcome to the Next Temperamental Era
Charles Schwab
Watch →