George Hotz | how do GPUs work? (noob) + paper reading (not noob) | tinycorp.myshopify.com
Key Takeaways
The video discusses how GPUs work, including their architecture and programming models, and covers topics such as parallel programming, Rust programming, and paper reading. It also touches on trust in society, demoralization, and erosion of trust, but the main focus is on GPU programming and related concepts.
Full Transcript
All right. Good morning, everybody. Welcome to Saturday in Hong Kong on Saturday. Um, oh, we could play on that laptop. No, that'd be a bad choice. Uh, see if, uh, wait for a few people to to get in here. How are How are you? How are you today? Today's Today's stream is gonna is going to be about you and it's going to be about education and teaching you things. Um, I'm not sure how possible education is. I You really have to ask that question, right? Everyone thinks that, oh, we're gonna like I don't know, maybe people don't even think this anymore. We're going to really get answers to that question, though. One of the things I like to think about is uh AI, similar to God, is going to judge you. Like, I think about it, right? I'm I'm applying right now for uh for Hong Kong residency and I have to put a resume together and oh this is just like like I got to put a resume together. I got to like like do like I got to like think oh wow remember that time I like went to college and like worked at Google and worked at Facebook and stuff like oh wow you know you got the normie you know that guy the guy the uh interview coder guy Roy Lee I love I love watching the people like this guy's destroying his future like dude what what world do you do you live in you know like oh no oh no he's he's he's rejecting stagnant tech companies and Columbia University. Like, wow. Oh, no. He's destroying his future. Like, what future are you talking about? Um, my advice to him is make sure you're always punching up. You know, you always got to make sure you're punching up because once you start punching down, you don't have the uh you don't have the uh the people's mandate anymore. So, you know, I always try in my life to punch up. And then once you get to the top, you got to stop punching. Like once you're the top, you got to stop punching. You got to be like, I'm the top. I don't punch anymore. We just try to build good things for many people. So, uh Oh, yeah. No. Oh my god. He's going to be so like like he's going to be so fine. What What are you even like talking about? Like like do do you understand the normie person who like graduates from Colombia and gets an interview at Amazon? Like I wouldn't hire them at Tiny Corp or uh probably, right? I wouldn't hire the statistical average of those people. I'm not sure I'd hire that guy either. I mean to be honest, I don't think he wants to work at a at a at a at a company, right? I think that there's definitely some uh some some some uh you know, questions there. Um, but I think if he really wanted to, I think he's definitely capable. Uh, you know, there's just the type of person. I mean, I'm kind of like that type of person, right? Could I really work at a company? I don't know. Maybe I'm unemployable, guys. I'm unemployable. What What am I going to do? What am I going to do? The system The system rejected [Music] me. What's happening with my utopia? I I did work at Twitter. That's all right. I was a fiveweek intern at Twitter. My utopia. You mean China? I don't think it's a utopia. I don't think utopia is real. Uh but I do think that Oh, I saw Oh, there was this guy. Uh let me find it over here. Um, someone posted on uh on Twitter that this is similar to my version of uh of Nobody Profits. Chinese subsidies uh create not destroy value. Um, and it's it's it's pretty good. Uh, like, do we really want tech billionaires or we want tech? In fact, mega cap valuations indicate something's gone seriously arai. What China wants from BYD and Jeno Solar should be affordable EVs and solar panels, not trillion dollar market cap stocks. Like, holy shit. Yes. Uh yeah, like like like make cars. Um you know, they talk about the global south a lot. Uh yeah, you know, so I I think that that it's it's nice to see that I'm not like way off in the weeds um thinking this stuff. And uh I read another thing last night. Maybe we can find it. It was on HackerNews. Um, it was some like Yimi crap, you know, like like the oh, we just need to build more housing kind of people. Uh, I don't know if we're going to find it. No. Oh, here we go. abundance isn't going to happen unless uh politicians are scared of the status quo. And there's a general fear of being a landlord. Tenants have a lot of legal rights and the risk of inviting someone on your property who could start squatting or doing drugs and not being able to evict them is beyond the pale for most families. What the hell happened? Like, and then all of these comments go on to talk about a legal balance between the landlord and the tenant. And something that I repeat all the time is that legal power is weak power. It's not strong. It's not strong power. You want strong power. A gun in your face is strong power. Saying, "I'm going to take you to court." And then after paying lawyers millions of dollars and having a judge and a trial by my jury of peers in three years, I can finally get a settlement which I then have to collect. Like that's such weak power. So it's interesting that the these things go that these things go back and forth, but that's the debate. And my argument is that the minute you've framed the debate as a question of the legal rights of the tenant versus the legal rights of the landlord, you've already accepted that you're in a medium trust society. Um, in a high trust society, these things aren't a problem. If someone starts squatting or doing drugs, their family will come and they'll be shamed. And, you know, it'll it'll you you'll deal with this by like, why would you do that? you know, he's a good Christian boy and he goes to my church and like you can make fun of that stuff, but that is what it means to be in a high trust society. Uh, you know, I was thinking about Hong Kong is a high trust society. Uh, and I think that's one of the main things that I really like about being here because say there was like a typhoon and you know the bottom floor flooded. I would be more open to letting random uh, you know, Chinese people who live downstairs and don't speak English stay with us uh, than I would if the same situation happened in San Diego. And it's it's it's sad that that's true, but it reflects the natures of the society. In fact, I'd probably let the person in in in San Diego stay, too. But I'd be a little bit more on my guard. Okay. Blah blah blah blah. Like here, I don't really know how things are dealt with, but things just generally aren't problems. Uh maybe one of the reasons is if you think about what the Christian morality says, the Christian morality says you should help the poor, you should help the beggar, uh you know, help the person who's down in their luck. And it's it's good morality, but that morality assumes that that's a very small percentage of the people. Once that percentage of the people starts to become the these people are saying like one in 10 of your tenants is going to become a squatter. Holy shit what a broken society. Um in Germany tenants have more rights and legal protections than landlords. Again, you can you can the minute you're arguing that I like there's a thing that I'm in favor of and that's the first thing that you think of, but then you realize the minute you're actually arguing that uh you've lost you've lost the plot. You you've you've you've accepted that you're not in a high trust society because in a high trust society, these things just aren't that big of a problem. Uh so I that's the big question and that's the big I do think that that to quote my blog post the demoralization is just beginning in America and it's that erosion of trust over the last 50 years. I wonder if there are polls on this. Like um USA poll, do people trust their neighbor? Okay. Here. I mean, this is this is here here's here's a here's a proxy for it. It doesn't have to be a neighbor, but like like and again, I didn't I didn't I didn't uh plan this beforehand, right? This is this is just my feeling. And then I Googled this. Um this is this is the closest question you can get at uh do you live in a in a in a high trust society or not. And you can see that even since the 70s when I I don't I'd be curious to see this going back to the 50s and the 40s. Um, yeah, I mean, you can see the the trend and I'd also like to see that continue and I suspect the continues on that same downwardly sloping line. Uh, it's not that I'm never going back to the US. I'm going back for the summer, but this is a major problem. And this this is the root of all of these other problems. And it's so obvious living here in America. I might, you know, start on some dumb crap about how tenants have way too many rights. Landlords should be able to immediately evict people if they don't pay. I like I agree with that. But again, the minute you as soon as you've gotten there, as soon as you've gotten to the point, I say the same thing about crypto. Crypto is a good crypto is the lowest trust society you could imagine. If society needs crypto like like the money like oh well we have a smart contract and we're enforcing this stuff like like you're just in such a it's a backs stop. Yeah, it can get worse. It can get but in a high trust society you don't need that kind of stuff. Um, no. And it's it's it's interesting how it how it rubs off on me how how I find myself wanting to participate in society in a high trust way. And I think that this is a human universal. If you put people in a high trust society, put one individual in a high trust society for the most part, they become more like their society. But unless this is fixed, it's over. Uh, let's see if we can find the same thing for [Music] China. What? They're asking people in other countries about whether China respects the person. What? What a joke. [Music] into USA again. I mean, we'll speculate for a minute. Uh, just imagine you were in the right like like again a media America's a medium trust society. I'm not worried in San Diego that all my neighbors are gonna gang up on me and kill me because I'm a libertarian. Uh I'm not a libertarian, but you know, people have people have misconceptions. Whereas like in in Cultural Revolution China, that shit might have actually happened. Uh so I imagine that when you compare China in the 40s and 50s to America in the 40s and 50s, America's way above China. Um but now it's changed. Uh, and if you haven't been to to East Asia, if you haven't experienced this, uh, yeah, I mean, I find that a lot of people in America are fed a lot of misinformation on this stuff, and until you fix this, you're never going to fix it. Oh, finally, finally, finally, we tweaked the law. We tweaked the law between the tenants and the landlord so there can finally be peace. Like dude, that's a dweeb. I want to like this like Dwight Shrew energy. Like like you just just No, that's not how it's going to be fixed. Nothing's going to be fixed like that. And as long as people are thinking like that. Yimi action. This doesn't work. Yeah. Let's make everybody scared. That's what everybody needs. Yeah. Let's make everybody scared. Um, like like Yeah. Okay. I think you get what I'm I think you all get what I'm saying. There's money and power and division. Yeah, you can argue that that's why this happened. But all right, welcome to the opening rant. And now we're going to go to subscribers only because now the opening rant is over because we are no longer in the opening. We are in the midame. We're in the midame. We developed our pieces nicely. Got good control of the center. Got a lot of Got a lot of squares covered that knight. That knight on F3 is looking real looking real flexible right now. The coffee we ran out of milk. We ran out of milk. Um, yeah, this is a good article. Uh, 5G here. I don't know what 5G is in America. I don't think it's that fast. Um, so let's start with the basics of what a GPU is. Maybe we'll start with the Nvidia CUDA. Have you guys ever programmed for GPU before? Let's just we'll do it again. So, have you guys ever programmed for for for a GPU before? Let's uh let's do a snap pole. And let's say can we do a pole? Oh, yeah. We got to do straw poll. I know how to do this now. All right. Straw pole. Have you programmed programmed a GPU at CUDA level? CUDA hip open CL level. Yes, a lot. Yes, a little. No. All right, let's see what we're starting with because this is a noob lesson. This is a noob lesson. We're gonna We're going to teach the noobs things. Okay. All right. There's the link. Metal. It's all the same. It's all the same crap. Oh god. Oh god. Really? Why do you even watch my streams? Uh, why do we only have 214 viewers? We're weak on viewers today. Wow. Wow, you guys. I don't know. Maybe maybe all the all the pros are sleeping. Why don't I have any viewers right now? Why am I only 214? Where's my viewers at? Okay, Friday night, late night, maybe. You know what I did last night? I played Magic the Gathering. Okay. Um, all right. We're going to have to start with some real noob shit then. GPU noob. Should we try to code in C? Does anyone want to code in C? Oh, let's code in Rust. Oh yo, I never code in Rust. I'm so bad at Rust. You can all make fun of me because I'm terrible at Rust. Okay. Um, all right. Let's see. Rust Open CL. Let's see how to do this. A Rust implementation of the Open CL3. All right. How do I use Rust? Oh, I need an LLM. Oh, how am I going to get an LLM? This is so annoying. All right. Well, we're stuck with this. Eventually, they're going to shut this down and then we're not going to have any LLM anymore. Okay. Um, how do I use Open CL from Rust? All right. So, we're going to we're going to bring myself down to the noob level because I don't code in Rust and I don't know Rust. Uh, and we're going to we're going to really start with some noob stuff. Okay. Create a pro Q. What? What? This is broken. Oh, this is the inside. This is the thinking. What? This is broken to use. Okay, we're going to need a cargo toml. Oh, it fixed itself. Okay. All right. How do I enit Rust project? Creating a new project. Okay. cargo new GPU noob. Oh, people are going to lose it because I'm because I'm because I'm writing Rust shit. All right. Now, let's not call it GPU noob. GPU noob will be the whole thing. We'll call it uh simple CL. Oh, that one created faster. Great. All right. Uh, okay. We got a hello world. Okay. Good, good, good, good. I I like the minimalness of Rust. All right. So, what do I do? Can I do cargo run? All right. Sweet. Let's add a dependency. Al 19. Is this called alle? It's called Open CL. on the seal crate. Is this really implementation of the Open CL API? How do I know which one's more popular? Aqual is here. Cl3 is here. What? That's Cl3. All right. Which one should I use? Do I want to use Aqual or open CL3? I'm going to make myself make my image a little smaller so we can see more [Music] code. All right. So, how do I judge what a good thing is? I don't know. Three sounds better than not three. should uh use a or open CL3. There is a more lower level binding beginner friendly Oh, well, no, but then there's CL3 which builds on top of this. Oh, this is based upon the Pro [Music] Q. You're new to Open CL and Rust. Oh, I've heard about Rusticle. Open CL implementation. That's just not going to work. All right, let's just go with this one. Let's go with the OCL crate. All right, so we include the crate in my cargo toml. Oh, look. I'm a Rust programmer. Do I have to use four spaces? Can I use two? Just one of these languages off the standard. Come on. Fine. Connect. That's fine. Okay. Good. this work. Okay, that's great. Okay, we got our nice code that we stole from an LLM and it works. Remove that mute. Oh, I see why it thinks it's mutable. Okay, that's good. Great. Work size. Global work size. What? Oh, I need some like Rust plugin or something. What did all Does crab show up? Well, it's kind of heavyweight, but okay. Uh, what about locals? I don't love this. Let's read the docs. Which one did we use? Used OC. That's all the types. Yeah, that's kind of cute. All right. Can I see some examples of this? This might be too high level. I'm not sure what it's abstract. It's going to work. No extra argument. What's d spatial dates? All right, cool. That works. Uh, what about local ID? Okay. local work size and the local work size might be set automatically. How do I override it? Ah, set default local work size. Okay, cool. Um, read the data back. Yeah, yeah, yeah. That's great. Yeah, it didn't work. Okay, that's fine. This number is a little big. Uh, let's just change that to like 26. I don't like all this this type stuff. How do I turn it off? Rust analyzer. Stop doing that. No, it keeps doing that. Okay. How do I turn that off? What is that? No, I don't want any of that crap. Thank god it's gone. Yeah, it's fine. No. No. Go away. Oh, fine. Keep on hoping we eat cake by the ocean. Oh, it turned itself back on assist. Okay, I want this off. Let's find settings. Is there settings? Settings. You guys get to see me be a noob. Is everyone excited about that? Rust settings. Extensions. Rust analyzer. Here we go. Uh signature info typing. Uh no no. Okay. I can change the settings. Show syntax tree settings. Inlay hints. off. Great. Oh, I see. So, it is a global setting. Thank you. Thank you, chat. I figured you guys would know this kind of stuff. Great. All right, let's go back in here. Darker run. Uh, let's print these things out. God, how do I print and rust? print the output in the C loop. Thanks LLM. Oh, that's not what I meant at all. That's not what I meant at all. All right, we should learn some Rust and how to print things. positional arguments. Got it. Um, I format this a little bit nicer. 56 is a lot. I don't know. 128. Okay. Getting an output from the kernel. This is kind of terrible. First of all, I print this like that. That of course that's wrong. You can pad numbers with extra zeros. Right. Justify with that. Okay, that's reasonable. Uh, now how do I print without the new line? Same as for the [Music] text. Great. Let's just like some rest. Yo, boys, we have vibe coding. Does anyone like this vibe coding? Does that work? No, probably not. Um, if I mod 16 [Music] 0. Who likes my vibe coding? Yo, we vibe coding. All right, great. We'll probably go back to Jersey 6. Jersey probably a nice number. No, 128 is fine. Okay, now we are Why is there a percent there? That's terrible. Y'all can make fun. Who's You all can make fun of my bad Rust skills. Uh, it's not actually mutable because we don't change that. where that can go. Okay, let's talk a little bit about what this is doing. Um, so get rid of that. We don't actually use those buffers. Um, don't actually use those buffers. Okay, so this is the source of the kernel that it's running in Open CL. Um this is a magical method called get global ID. Open seal is a little bit weird about this because it doesn't exactly it doesn't exactly uh you can think about GPUs really as well. We're only going to use one global ID here uh and one local ID because that's all that there really are. So you can think about GPUs as G and uh G cores with L threads. So this is your in Open CL this is called a group. Uh this is called a local uh in CUDA they have different names. Let me just see this code in in tiny grad. Forget what I forgot what they all they all have like slightly different names for this grab and it's like uh shiny grad. No, no, don't. No. In a new window, please. So here's tiny look in C style and see what they each call them. So you see in open CL they're called so like the functions are like get group ID and get get local ID. We'll get into what these are in a minute. So this is Open CL. Um, CUDA calls them the block idx and the thread idx. It's actually a much better name. What does hip call them? Does hip have a stupid name for him too? Um, hip calls them group and local. Same as uh cuda hip. That's like the real hip. What does metal call them? Metal calls them thread group position in grid and thread position in thread group. Okay, but they're all basically the same thing. And you can think about it as six cores and L uh so G cores and L threads. Okay, so um I don't really know why this has a dim here. What happens if I get rid of that? What is what complains? ProQ has not any dimensions specified. Okay, let's specify some dimensions. So here we should be able to say global work size. Actually, is it global work size? There's three dimensions, but you never have to worry about the three. What? I specified the dimensions. Oh, stupid Rust. This is actually stupid OpenC. I shouldn't blame Rust. It's not Rust's fault. Okay. So the yeah the stupidity of uh Open CL is it specifies this thing called dims and dims equals G * L. Okay. So in this example, G is 128 and then L is specified down here as 1. So if we want to see what those things actually are, we can go here and we can get the group ID of zero. And you'll see that right now we're running on 12 uh8 cores and one thread each core. So we can run 128 cores, one thread each. Now if I change this local work size to be two, how many cores and how many threads are we going to run? I'm going to run 64 cores with two threads each. So let's see if we actually confirm that we're setting the group ID there to the output C. So what we expect to see is 0 0 1 1 22. But nope, that's not what we saw because for some reason it didn't actually listen to this local work size. Ah, no, it's dumber than that. So, you see what we did there? We didn't set the um let's make this get global ID zero for the actual output of the thing. There you go. Right. So, that's a global ID and that's a group ID. So now we're set to the group ID. So your global ID is your is your uh you want to take your so so you can think about it as like G and L. You can think about it as G times numbum threads plus L how it's actually stored. Okay. So now we can put it on four course. Let's put it on four cores, four threads. So you see that's thread. All right, we're good. Uh now if I you see I can do get local ID here and we're going to see 0 1 2 3 0 1 2 3 0 1 2 3 0 1 2 3. Okay, great. Cool. Is this making sense to everybody? Is this is this good? Is this good new lesson? If I unen if I enable nonsubscriber chat, are you guys gonna be terrible or are you guys gonna be like, "Wow, he's actually educating us and teaching us something because we're going to we're going to really all these lessons build on each other." So, I hope you're paying attention. You're like, "Oh, this is simple. Oh, I understand this." Okay, great. You understand this? That's great. Okay, here's a question. What if I make dims really, really, really large? How does it run? The GPU doesn't have this many uh cores. If the GPU doesn't have this many cores, how is it going to run? And the answer is, think about the G's. Each G is kind of like a packet of work. And it's uh if you're familiar with like Grand Central Dispatch on Mac, it's like the same basic idea. All the GPU cores are sitting and waiting for um a chunk of global data to become available and then they'll run them on the threads. So let's get some real numbers here about how many uh cores GPUs have. So how many cores does 4090 have? So Nvidia calls their cores streaming multipprocessors. Um on Nvidia cores are SM on AMD cores are uh compute units streaming multi-processors compute units My chat keeps breaking. I don't know why this breaks all the time. George looking like Steve Jobs today. Okay, now I think I'm caught up. Okay, course multiprocessors on AMD cores or compute units. Um, so let's get into some stuff. How many streaming multipprocessors does a 490 have? That's a 3000. Uh, the actual name for the chip is 812. There you go. Oh, that's pretty nice. 8102 right here. So, each of these GPCs, one, two, three, [Music] four. Let's just check with a P. So 8102 has 144 streaming multipprocessors. 4090 uh and actually the 4090 might even scale that down a little bit. I think this is just the full chip. So AD102 4090 let's do it. nothing disabled. It has 144 SM and it has this many CUDA corores. So how remember a CUDA core is like a th is like a thread. So [Music] there's 128 threads with 128 threads each. Well, I mean, this is education for everybody. So, this is I thought there was going to be 32, but maybe not. Okay. So, then the uh 700 XTX is called Mavi 31. So, Navi 31 has 96 compute [Music] units, 96 CUS with how many threads each? I think it's this. Well, so the cool thing about this is there's actually a PDF that shows you what a compute unit actually is. That's a workg groupoup processor. Don't you love all these names? So each workg groupoup processor has two compute units. Then each compute unit looks like it has 64 threads. No, maybe not. Actually, so dual compute unit 8 in each. 8 * 6 is 48. So it does have 96 compute units. That's a workg groupoup processor. And each workg groupoup processor [Music] has three sims. So it does seem like they have 64 threads each. All right. This doesn't seem right because these chips do have similar uh power but no maybe that's right. Well, we'll we'll we'll get we'll we'll dive in deep and we'll understand this. Um so yeah, each workg groupoup processor is these these these sims are are kind of like the the well each of these sims is 32 threads and that introduces the next concept on a GPU which is called a warp. Okay, GPUs have warps. Warps are groups of threads. Um, and all modern GPUs have [Music] them as 310. Okay. So, um, are you guys familiar with what SIMD is? Put my camera to the left, please. Who knows what SIMD is? Who can tell me? I don't know why this chat breaks. Is that live? The web saga gets disconnected or something. Um, okay. So, SIMD stands for single instruction multiple data. Uh, I don't know why chat's broken. Stream's working though, right? Um, CID stands for single instruction multiple data. So, maybe a simpler way to think about 70 is just to think in terms of vector registers. So if you think about uh a register on a GPU, let's say the register is 32 threads. Each thread can fit a single float. So if you have a, you know, it's just like it's like a vector of like float 32. Um how many bits is that? Well, that's going to be 32 * 4 bytes * 8. So that's 1024 bits. Um and then if I do something like C= A + B uh on vector registers, this is a single add instruction on 32 pieces of data. So the GPU programming model is slightly different from this. So if you ever use SIMD, this is how SIMD works. You can use this called like AVX or Neon or whatever you want to call it. GPUs use single instruction multiple thread. The math is pretty similar, but on a GPU, you never see it do float 32. Uh on a GPU, you'll just declare float. So notice how I just declared float here. So this or I can just declare int here, right? I can say like int a equals get global ID. Um, and then a is here. So you don't see it even though this is running on on four threads. And let's let's increase it to the natural number of 32. So it's a single warp. So now it's running on a single warp. So there's going to be four warps there. Um, single instruction multiple thread. So even though you don't uh you don't see that this is a vector register it just declares it as int vector register. What makes sim a lot easier to program for than sim is loads and stores. Loads stores are different. So the problem with sim is let's say I want to get some data into a. If a is a 32-bit thing and I do a subi, what that's going to do is load the contiguous uh 32 things. Um, load stores are implicit scatter gather because on a GPU I've declared it like this. I can happily do something like that and it will get them. Whereas on SIMD it's explicit. That's it. That's the difference. Uh GPUs like to hide this fact from you that what they really are is one24bit wide uh SIMD machines. But they really are 1024-bit wide SIMD machines, but they have this special thing called memory coalescing. Um, so you'll see that like GPUs can do all this stuff for you fast even when accesses are not aligned. They'll do all of that for you. Okay, great. Does everyone understand the difference between SIMD and SIMT? You only declare float behind the scenes. It's float 32. Okay, so you might think like why do I care? And for a program like this, the answer is you don't care. But let's see if we can get in a little bit and make the programs a little bit more complicated uh to the point that you might care. So, let's put a loop in here and have to deal with timing shit and rust. Um, Oh, perfect. Yes. Oh, that's pretty nice. How did Russ know to format it like that? That's very fast. So, let's add a stupid loop in here. Um let's we have to make it not get optimized out. Compilers are very clever. Uh also this is kind of wrong here. We'll have we'll put it after the read. Let's make sure nothing's getting optimized out. Okay, maybe it's not getting optimized out. It's probably getting optimized out. Let's think of something that's too stupid to optimize. [Music] Uh okay, good. That didn't get optimized. Doesn't know how to multiply by two. Okay, great. So, that takes 37 milliseconds here. That's probably good. Okay, 163. Okay, so what's going to happen if I only run this on one thread? Takes the same amount of time. No idea why that is. Oh, probably because globals isn't big enough. I'm going to have to make globals bigger. CNC data. stuff isn't used anymore. Not used. Not used. Let's make dims. Find a number that matters. Uh 32768. Actually, that's still too small. Oh, yeah. All right. That's chilling. That's chilling. All right. Two seconds. Oh, faster. All right. 220 seconds. Great. Uh, so now if I only tell it to run on one thread. Yeah. Should be about 32 times slower. So, is that 32? It's 229* 32. Yeah. Okay. It's about 32 times slower. Okay. Okay. Now, if I run it on two [Music] threads, twice as fast, right? This make sense? So, remember that uh this is not going to be your your global uh your global threads. This gets divided by this. So, like how many threads are we are we are we putting it on? Go to four threads faster, eight threads. Um, I also want to time this. I also want to change my timing a little bit. I don't want to include the read in the timing. So, I should be able to do something like ProQ finish. Yeah. And we'll do timing there. This result may be an error variant. Okay. Right. Uh so 8 16 we'll watch it still get twice as fast. Should be about 420 milliseconds. uh 32. But something interesting should happen at 64. At 64, I predict it will no longer get any faster. Yes. Does everyone understand why they tell you this stuff about how these things have multiple threads, but I never think about it like that. uh we we'll get to well I don't really understand why that is but you can really think about GPUs operating uh 32 wide oh I know why never mind I understand why okay you can think about all GPUs operating basically GPUs are multi-core processors with 32 threads um so we're on a Mac let's actually figure out what we can what we can do. Uh it's there's a lot less publicly known about uh this stuff, but you see why putting it to 64 there no longer made it any faster because GPUs only have 32 threads. So what you're doing if you're going to 64 is you're going to run this chunk and then you're going to run this chunk. It's just what happens. So if I go to 16 and then I think 128's not going to give us any more speed either. And then eventually, let's keep going up here. 256. Let's go to one 512. Ah, 512 doesn't work. 256 is max threats. Now, if you're like, "Oh, they're running sequentially. Oh, why does it even matter if it's uh actually is 256 the max?" 27. Okay. No. Good, good, good, good. We have we have like some plausible numbers here to work with. Okay, so let's see what we can find out about Apple GPUs. M3 Max GPU. So, it's a 40 core [Music] GPU, but these cores can do a lot more. Um, well, we can try to figure this out by profiling if we don't have another way to get it. Okay, here we go. So, the M3 Max, I have the big M3 Max variant. It has six uh six execution units, I assume that stands for. Yes, execution units units. And then it has this many threads. Let's see what how many they're saying is per eight. That doesn't make sense. It doesn't make that much sense, but okay. I think we can figure this out if we we can figure out when it stop we can change the global size here and figure out when it stops getting faster. So, this says there's let's just try that to start with. So, I set the local size to 32. So, that's going to put it on 320 cores. Um, I don't think execution units are exactly cor. We'll figure this out. Okay. So, that's 320 cores. How much slower is it when I go to 600? Oh god, it's a little slower. It tells me nothing. We'll start with a one. That's fine. So, one takes 36 milliseconds. So, 36 milliseconds is my minimum. 36 milliseconds. uh this kernel takes 36 milliseconds to execute. So think about it. You got threads. Uh actually just go to one here. Oh well that didn't work. There's a working size. Yeah. Set that to one as well. There. So, this kernel takes 31 milliseconds to execute. That's just it. Takes 31 milliseconds to run this stupid loop. Uh, we can try to figure out why that is in a minute, but well, let's see if that's remotely plausible. Okay, so it's um [Music] 1,00 so 26. Uh, it's that like doesn't seem that plausible. GPU is like slow. Only a million. This thing should be running at gigahertz. I understand why that's so slow. Oh, there's like a base. I see. Okay. [Music] Um, yeah, it's not that great. [Music] Uh there's like a there's like a base thing that's just uh Oh, well, you know what? Here's something we can do. Let's not look at the first run of it. run it a few times. All right. A for loop in Rust [Music] for warm up in like that. Oh, okay. It's just slow. It's fine. Let's just do we can get that 32. Great. Okay. 14 milliseconds there. There we go. Now, now we're seeing Now we're seeing good scaling. So, it's 14 milliseconds to run. The first one's just slower, but now you see when I add one more zero, it adds one more zero to the thing. So, the kernel takes 15 milliseconds to execute. Um, that's one E6. It's in 15ms. So that's let's find something 26. So that's something like 66 uh million per second. What's the clock speed of GPU? boost up to 1,600 something like 24 instructions cycles per seems somewhat plausible. Cool. Okay. So, one of these runs in 15 milliseconds. If we do a uh global size of one. Oh, that's going to be really slow. Okay, so 256 didn't get any slower, which means that the GPU has 256 cores. Remember, GPUs have cores and threats. So, it's easier if you just like look at one. Uh, I found this thing high yield. This guy goes into the chip. That's cool. So, it's it's it's helpful like everyone needs to really of the entire chip, the compute parts, everyone needs to really understand what's going on here. Uh before we can move on. All right. So, this is the chip. Each of these compute units is a core. And then inside the compute unit, we can zoom in and we get a thread. So in each of these cores, each of these threads is 32. So each of these is like a a little computer with 32 threads. See how everything's 32? So each of these compute unit pairs has like four cores inside of it. And those are full cores. So that's 8 * 6 is 48 is 109 * 4 again is 192. So there's 192 cores each with 32 threads. So 256 is still fast. Let's find when this starts to slow down and we can figure out how many cores this thing has. Okay, so that's loaded down. That takes twice as long. That means it's taking like some cores have to run, too. So, let's see if it's really 640. 640 is a little slow. 600's a little slow. That's still fast. 513. 513's fine. 520. You know, I used to say hacking. I was just binary searching by hand 540 50. Ah, that got slower. Should we like make plots and shit? We should make plots. What's that? Who's ready to make plots? Oh, let's make plots. God. All right. All right. You guys are really You guys are really testing my rust skills here. Should we use rust to make the plots? I don't even know how to write a for loop in Rust for num cores in remember this here is fixing it to be only on one thread. Guess where do I set the dims? Wow, I set the dims all the way there. Can I change the dims? I want to update the dims. There we go. [Music] Data length exceeds buffer length. [Music] What? What? Where does it tell the size of the buffer? The default dimensions will be used. Oh no. Not cool. Not cool, bro. Why would What? That doesn't even make sense. buffer builder doesn't even make any sense. Why would you use that? Well, I'm go make me uh buffer. Shit. All right, here we go. Okay, here we go. Buffer builder. All right, let CB buffer equals Did I have to import the freaking buffer? Buffer builder Q flags len build question mark. Oh, wait. What is the question mark though? I think I just want to use that question mark in some places. Why did you ignore that? Okay, good. Great. Thank you. That'll make sense. [Music] Oh, I broke it. Oh, I didn't break it. Oh, no. Oh, I see. Just multiply zero by zero a lot of times. All right. So, how do I like lie to this and get it to like not Okay. There you go. Yeah, I used it. You see, I used it. Yeah, it's legit. All right, cool. Um, yeah. 32. No, but it didn't listen. A bish don't listen. Oh, I know what I can do. Oh, I just didn't do better. The dims don't really matter. There we go. Okay, now it's only running at 32. Great. All right. for test scores in 0.124. Did that work? Cool. So, we're going to see a jump when we get to the number of course. All right, we're going to make a graph. Can you make graph and rust? Is that doable? Uh, we can also how do I like that? All right. Step and rust loop. Step in loop. Step five. [Music] I want to make graph graph range step inclusive. Yeah, good. They got rid of inclusive. That's good. That's good. Rust is moving in the right direction. All right. uh make graph in rust. I don't know why the chat keeps disconnecting me. Yeah. You like my you like my new lesson? Is everyone happy? I'm teaching you something hopefully. Have you guys never coded on a GPU before? All right. Bit map back end using the pet graph recommended. No, I want a chart. Use plotters. Plotters. Is this good? Let's try plotters. Plotters. Oh god, I need to make an array. I'm going to do this guys. Rust is impossible. Can I put the question mark there? I want to bang there. Can I bang? No. Question mark. Does that make you happy, Rust? Rust is never happy. Oh, good. We can remove that meat. Don't meat. All right, bitch. I want to plot. Let me go back. V. Yeah, V. That's a good name for a ve dot. No, no, no. Don't put it in a box. Oh, do I need to tell it what it's a ve of? What's this a ve of? Man, I wish I had like like the type hints would show up, you know? Why does my chat not work anymore? Why? What? I hate this. Why Why is that doing that? I Why can't I make a ve? um make that vec, but I can append to it with a tpples the plotters. Yes. Oh yes. I want mutable vec I32 I32 except no I want float F32. Oh yeah yeah yeah yeah. Let me points equals vec new. Yeah yeah yeah. There we go. There we go. Yeah. Yeah. Push those points. Yeah. Now we're talking. Oh man. I'm a Rust programmer. All right. It's like the LLM is a Rust programmer. I'm just sure to criticize it. Why a error? Why is it Oh. Oh. Fine. Fine. You got a box. You happy there's a box. I got I bought you a box. Oh. No, no, don't do that. Becca, do I have to import block? Hey, hey, yo, I'm getting errors. Yo, yo, I'm getting errors. Yeah, please please like just like don't like fix them. Yeah, we're vibe coding, guys. We're vibe coding, right? Yeah, you did a bad job. So, you got to parse all that crap. Oh, yeah. Okay. Well, the user's main function returns result. Oh, maybe it should return result, aqual error. Oh, great. Yeah, but what? No, no, no. You're using generic error. Yeah, yeah, that's right. That's right. Yeah, yeah, yeah. I got a box. All right, good. I put it in a box. All right. Is that good? Is everybody happy now? I put the error in a box. All right. Is that good? No. Oh, Russy explain. All right. Here. Oh, no. Not good. No. No. Why is it still trying to do a error? Let's import that one first. Yeah. Yeah, that work. No. Yeah, that's right. Great. They have an alias. I don't know the alias. I put the error in a box. Who knows what I did? Oh. Oh, result. Oh, I see. What a scam. Oh, a huge scam. Okay. Why do you keep using result? No result. Just error. results imported. What? Oh crap. Expected duration found I32. Well, we're not going to sit here and wait for that every time. So, a lapse, unfortunately, is a duration. I go like two float. Uh what's a word as millies? A millia here, a millia there. That error is backwards. U238. Okay. as I 32. Okay, just draw my plot. Oh yeah, look at that piece of shit plot. Yeah. Yeah, that's exactly what I wanted. Yeah. Wow, you really read my mind. Rust actually kind of like rust. No, why can't I put that in there? Oh, no, no, no. Don't get your those things 14. But no, no, no, no. Come on. I want like a float. No. No. Come on. Come on. I want like a float. No. Don't do that. Don't do me like this. Come on. All right. Fine. Fine. Fine. That's That's good. That's good. That's good. All right. You happy? Everyone's happy. All right. So, where my where my plot at? All right. Well, what was the plot? Oh, did I not like I got to like tell her like I should do something? Did I not copy and paste something? Use the collected points. What? Save it. There's no line in the pinch. I'll do how I get line. vibe coding. I love LLMs. Oh, is this the problem? Oh, that's the problem. Maybe. Why would it ever do that? Oh, the range. Oh, okay. Thanks thingy. I don't know. Let's go to like a 100,000 here. And we'll go to like I don't know. One, two, four here. Yeah. Base line plot. All right. 100,000 might be a little overkill. Let's go to 50,000. And let's go to 1024 here. Yeah. Oh, check out baseline plot now. Yeah. All right. Does everyone understand how many cores it has? I don't really understand why it's doing that, but we see a hard change over at at 640, I think. So, there are 640 cores. And then theuler is just like kind of mediocre. All right. What's margin? What's this? the size of the four margins of the chart. Why are there so many lines? All right. But uh so it is 640 where it hard switches. No, it's not 640. Oh, yeah, it is. Okay, good. So, it does have 640 course. Okay, good. Everybody's happy. Just the schedule is kind of shitty. So, it really does have 640 course. I don't know why it only has that many ALUS. Now I expect Okay. Can I do like this? Of course, that's not going to work. Syntax. cannot use points in this scope. That's actually a reasonable error. Orange is not a color. Yellow. Yellow is a color. Yellow. No semicolon parenthesis. You be happy now? Good. You be happy. Good. Everybody's happy. H. Oh, whatever. We're going to start at 64. All right. So, now we're just What? Why is that group size? That don't make sense. Oh, because it's not like a multiple. That shouldn't really be required, but quick fix. Yeah. Sometimes you want max. Yeah, look at that. Yeah, you know, it's going to be mad. You imported min and you didn't use it. Okay. I went to all that effort to import min for you and you didn't even use it. Yeah. Okay. Baseline plot. All right. Uh, well, that's what happens when you put in threads. I don't know. Let's go to two there. Um, let's make that a little faster. Well, 640. That's a divisible by 16, so that's good. Uh, let's go a little bigger here. Oh, no, no, no, no, no. We made a mistake. We made a terrible mistake. The charts, man. Impressive. Yeah, right. Don't you love rust? Rust kind of great. Is rust going to be like my new thing? H That's the ugliest yellow I've ever seen. Seion. Oh, no, no, no, no. Magenta. Magenta. That's the perfect color. I don't know. Uh, let's try times locals there. That's not really what I want. Let's just step by 64. And let's go to something crazy there. Okay. So if we have 640 threads 32. Oh, it's looking slow already. Look at how slow it looks. Oh, it's gruelingly slow. Oh. Oh. I'm not going to sit here and wait while that happens. What? What? Come on, be happy. Okay. Um All right. So, yeah. Yeah, this is kind of notice how notice how beyond here it doesn't uh it doesn't matter anymore because this is how many threads the GPU really has. [Music] So measured We measured 640 EUS and with 32 threads each We can add some more colors in here for I don't know. Actually, you know what? Could we just make them all the same color? Does everyone have a good understanding of what this is going to look like? So you can see um should actually do times 2 times locals with min Well, now I need min. Oh, let me just import min. If it tells me that I could put the comps together. No, it didn't tell me that. All right. So, does this make sense? That's with one thread. 2 4 8 16. But then 32 and 64 are the same because there really only are 32 threads on the GPU. So, all it can do is run this thing twice. Doesn't make sense to you. Yes. All right. You get it now. That's good. That's good. Great. So, this is the minimum amount of time that this kernel can take. That's just this kernel. If you just run this kernel single threaded on one thing, then there's a question of how many threads you want to run it on and how many cores you want to run it on. So it'll dispatch to all the cores. And this means it has to run in series. So this mean has to run serially. So it has to run it twice. All right. Does everyone understand baseline plot? Because we're going to we're going to get rid of baseline plot in a minute. Let's uh let's actually commit this stuff. Might be useful for somebody. Um, no commits yet. What was this? Some like Rust. Rust automatically creates a git shit for me. I don't know how I feel about that. I really don't know how I feel about that. That's a little upsetting to me. Did it name it Maine? If it named it Maine, I'm going to lose it. Oh, good. It named it master. All right, we're good. We're good. Everything's fine in the world. Those look like good files. All right, I'll let me create a thing for you quickly because we can also start to run this on tiny boxes if we want. All right. Noob lessons from stream about how GPUs work. Let's add a remote and let's push that shit. Oh no. Come on. [Music] Yay. All right, you guys can go to the repo now. Okay. So, we made baseline plot. Everybody understands it and we're good. Uh, what else do we want to check? I mean, I think this kind of explains this kind of explains the GPU execution model to everybody. I don't know why it says that that's only that that's the number of AUS only 500 doesn't seem right. app calls a GPU core seems to be the same as what Nvidia calls an SM. each core. It's interesting. I don't really understand It's mad cuz it moved. It's fine. All right, let's let nonsubscribers talk for a little bit. Let's talk about what we don't understand about GPUs. Okay, only serious questions. Banhammer is going to go hard. I think that like my chat's broken. You're asking me to ban you. Okay, done. How they settle on 32? Uh I don't know. Nvidia did 32 and everyone just kind of copied them. Is this Rust code vectorzed? So this chunk here is what's running is what's being dispatched to all the uh to all the GPU execution units. So you can see that the dispatch happens here and then I wait for them all to finish here. then like hopefully you understood what I said about the course and the threads and stuff. Time zone am I in? I live in Hong Kong. Uh yeah, G. Well, sort of. So, I think we talked about that. It's not really SIMD. So, it's SIM T. So, G the the GPU that I'm working with here is basically 640 cores with 32 threads each. It's a little bit more complicated than that, but uh does AMD still have 32 and 64 thread modes? It actually does, and that came up this morning. Um there I said, so AMD has the best documentation uh for like all this stuff they'll talk about. So uh AMD doesn't call them warps, they call them waves. Okay, waves. Um, so you can get 32 and 64, but all the code I've seen only outputs 32, but there's apparently something outputting 64 now. Good website to understand GPU nomen glitch. Let's see what we got. Am I still holding AMD stalk? Yeah, I'm holding for a long time. Oh, interesting. Who's modal? Um, why is it green program GPUs? Yeah. Okay, cool. Oh, it's pretty good. Yeah, this is good. Good. Good call. Yeah. So I mean then we can get into like latency hiding and how that stuff works. If you think about a CPU will context switch in order to hide latency from like a network card uh or disk. A GPU will context switch in order to hide latency from memory. Google's over. Google's dead. Um, so I've been working a lot on the DSP on the uh on this DSP on the Qualcomm DSP. So I've been I've been reading this manual a lot. Um, you get into like how the DSP works and the DSP the DSP is a SIMD machine, not a SIM T machine. So, all the memory accesses are annoying. And then there's like very specific patterns you have to use it in to be fast, but you can see the uh Qualcomm hardware channel. I've been posting all the updates. We finally been making it like fast. Uh and then comma can switch back to the DSP for um the driver monitoring model. But yeah, like when you look at these like look at that instruction, like look at what that's doing. So it's doing it uh for uh here we can see the vectorzed version down a bit further too. So like this is the vectorzed version of that. So you can see it's doing and this is this is not the whole thing. So it's actually 128 bytes uh which is 124 bits. It's the same,24 bits. It's the same it's the same widt
Original Description
Date of the stream 29 Mar 2025.
from $999 buy https://comma.ai/shop/comma-3x & best ADAS system in the world https://openpilot.comma.ai & - https://tinycorp.myshopify.com
Source:
- https://docs.tinygrad.org
- https://github.com/tinygrad/tinygrad
Follow for notifications:
- https://twitch.tv/georgehotz
Support George:
- https://twitch.tv/subs/georgehotz
Order tinybox:
- https://tinycorp.myshopify.com
Chapters:
TBD
Official George Hotz communication channels:
- https://geohot.com
- https://twitter.com/realGeorgeHotz
- https://instagram.com/georgehotz
- https://tinygrad.org
- https://geohot.github.io/blog
- https://github.com/geohot
We archive George Hotz and comma.ai videos for fun.
Follow for notifications:
- https://twitter.com/geohotarchive
Thank you for reading and using the SHOW MORE button.
We hope you enjoy watching George's videos as much as we do.
See you at the next video.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from george hotz archive · george hotz archive · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
comma ai Driving to self racing cars with openpilot
george hotz archive
comma ai Still driving
george hotz archive
comma ai was live
george hotz archive
comma ai Going home
george hotz archive
comma ai We go to the airport
george hotz archive
comma ai Reversing Prius with cabana + panda telethon!
george hotz archive
comma ai panda manufacturing!
george hotz archive
comma ai Self driving to Best Buy
george hotz archive
comma ai shilling for giraffe!
george hotz archive
comma ai Toyota Prius Driving!!!
george hotz archive
comma ai Late night civic driving
george hotz archive
comma ai Toyota giraffe shilling
george hotz archive
comma ai Live car hacking with panda this time or bust!
george hotz archive
comma ai Product launch question time
george hotz archive
comma ai Driving with the RAV4, launching Tuesday!
george hotz archive
comma ai giraffe ship o' clock
george hotz archive
comma ai openpilot 0.3.9
george hotz archive
comma ai EON assembly!
george hotz archive
comma ai Going through the GM investor deck
george hotz archive
comma ai I love my EON
george hotz archive
comma ai RAV4 driving
george hotz archive
comma ai Shilling at the holiday party
george hotz archive
comma ai EON shipping party
george hotz archive
comma ai EON unboxing!
george hotz archive
comma ai The very straight roads of Nevada
george hotz archive
comma ai Starting our trip with openpilot 0.4
george hotz archive
comma ai Little EON on the prairie
george hotz archive
comma ai The urban sprawl of Colorado
george hotz archive
comma ai Onward to Omaha
george hotz archive
comma ai nothing, nowhere
george hotz archive
comma ai shop.comma.ai Buy things!!!
george hotz archive
comma ai The youth are woke
george hotz archive
comma ai Photo shoot!
george hotz archive
comma ai Product announcements are LIT!
george hotz archive
comma ai Breaking down hype of CES
george hotz archive
comma ai Salt Lakes Everywhere!
george hotz archive
comma ai This is the last one
george hotz archive
comma ai Corolla port o’clock!
george hotz archive
comma ai Presentation where it’s like you are in Omaha with us
george hotz archive
comma ai Asking the scopies the banned question
george hotz archive
comma ai Driving in the Corolla!
george hotz archive
comma ai We got new products! shop.comma.ai
george hotz archive
comma ai Sunday w scopies!
george hotz archive
comma ai Our first Lexus, the Lexus RX!
george hotz archive
comma ai Scopie saturday!
george hotz archive
comma ai Panda!
george hotz archive
comma ai Scopie Sunday! *NOT CLICKBAIT*
george hotz archive
comma ai comma Tree!
george hotz archive
comma ai Scopie Saturday
george hotz archive
comma ai Ok scopie Friday
george hotz archive
comma ai comma pedal!
george hotz archive
comma ai okay this time comma pedal!
george hotz archive
comma ai Why aren’t car companies good
george hotz archive
comma ai How can driving be better
george hotz archive
comma ai Scopie Sunday
george hotz archive
comma ai comma got a new car!
george hotz archive
comma ai Mapping Sunday!
george hotz archive
comma ai Let’s go buy a car
george hotz archive
comma ai Ok I take back all the bad things I said about Ford
george hotz archive
comma ai comma smays are in stock!
george hotz archive
More on: Reading ML Papers
View skill →Related Reads
📰
📰
📰
📰
A lightweight workflow for keeping up with AI conference papers
Dev.to · Daniel
Why CitedEvidence Believes Great Researchers Read Less Than You Think
Medium · AI
How to Write a Literature Review That Actually Argues Something
Medium · Machine Learning
I Built a Personal Paper Engine to Stop Losing Research Papers
Dev.to · Ethan
🎓
Tutor Explanation
DeepCamp AI