Will LLMs Get Us To AGI?

a16z · Beginner ·🔢 Mathematical Foundations ·9mo ago

Key Takeaways

The video discusses the potential of LLMs to achieve Artificial General Intelligence (AGI) with Columbia CS professor Vishal Misra, covering topics such as chain-of-thought reasoning, token prediction, and the limitations of LLMs in recursively self-improving.

Full Transcript

Any LLM that was trained on pre-1915 physics would never have come up with a theory of relativity. Einstein had to sort of reject the Newtonian physics and come up with this space-time continuum. He completely rewrote the rules. AGI will be when you are able to create new science, new results, new math. When an AGI comes up with a theory of relativity, it has to go beyond what it has been trained on to come up with new paradigms, new science. That's my definition of AGI. Martine, um yeah, I know you wanted to have Vishal on. What what do you find so remarkable about him and his contributions that that inspired this? >> Vishal and I actually have very similar backgrounds. We both come from networking. He's a much more accomplished networking guy than I am, but That's a high bar given you your accomplishments in the field. But but what we come from and so we we actually view the world in an information-theoretic way. It is actually part of networking. Um and you know, with all this AI stuff, there's so much work trying to create models that can help us understand how these LLMs work. And in my experience over the last few years, the ones that have most impacted my understanding and I think have been the most predictive are the ones that Vishal has come up with. Um he did a previous one that we're going to talk about um called Matrix, is it? Uh Beyond the Black Black Box, but yeah, the Matrix >> Box. So actually if yeah, you know, we should put this in the notes for this, but like the single best talk I've ever seen on trying to understand how LLMs work is one that Vishal did at MIT, which uh um Hari Balakrishnan pointed me to and I watched that. So So he did that work and then he's doing more recent work that's actually trying to scope out not only how LLMs reason, but like it has some reflexes on humans reason, too. And so I just think he's doing some of the more profound work in trying to understand and come up with models, formal models for how LLMs reason. Which is on that note, you said his most recent work helped you change how how how humans think. Why don't you flush that out a little bit? How did it sort of Well, okay, so can I can I just try to take a rough sketch at it and then you just tell me how how how how wrong I am? >> I am. You know, you're trying to describe how LLMs work and one thing that you found is that they reduce a very very complex multidimensional space into basically a geometric manifold that's a reduced state space. So it's a reduced degrees of freedom, but you can actually predict where in the manifold the reasoning can move to. Roughly. So so So you've reduced the dimensionality of the problem to a geometric manifold and then you can actually formally specify kind of how far you can reason within that that manifold. So it it it it and the articulation is that we or one of the intuitions is that we as humans do the same thing is we take this very complex heavy-tailed stochastic universe and we reduce it to kind of this geometric manifold and then when we reason, we just move along that manifold. Yeah, I think you captured it accurately. That's that's kind of the spirit of the work, yeah. >> Well, wait, can I just hear it in your words because, you know, I'm this lay I'm a VC, so You're you're a VC with an H-index of what, 60? Something Uh yeah, so so you know, ultimately what all these LLMs are doing, whether you know, the early LLMs or the LLMs that we have today with uh you know, all sorts of post training RLHF, whatever you do, at the end of the day what they do is they create a distribution for the next token. Right? So given a prompt these LLMs create a distribution for the next token or the next word and then they pick uh something from that distribution using some some kind of algorithm to predict the next token, pick it, and then keep going. Now what happens uh because of the way we train these LLMs, the architecture of the transformers, and the loss function, you know, the the way you put it is right. It sort of reduces the world into these Bayesian manifolds. Yeah. And as long as the LLM is going uh in it's sort of traversing through these manifolds it is confident. And it can produce something which is which makes sense. The the moment it sort of veers away from the manifold, then it starts hallucinating and starts spouting nonsense. Confident nonsense, but nonsense. Yeah. So so so it creates these manifolds and the trick is, you know, we the distribution that is produced you can measure the entropy of the distribution. It entropy the way Shannon described >> Shannon Shannon entropy. Shannon entropy, yeah, not not thermodynamic entropy. So uh So so suppose you have a vocabulary of let's say 50,000 different tokens and you have a distribution next token distribution over these 50,000 tokens. So let's say the cat sat on the right? If that is a prompt, then the distribution will have a high probability for mat Yeah. or hat or table and a very low probability of let's say ship or whale or something like that, right? Yeah. So because of the way it's trained, it it has these distributions. Now the distributions can be low entropy or high entropy. Yeah. A high entropy distribution means that there are many different ways that the LLM can go Yeah. with a high enough probability for all those paths. Yeah. Low entropy means that there are only a small set of choices for the next token. And the prompts also you can categorize into two kinds of prompts. One prompt is uh is you can say uh high information entropy Yeah. and one prompt is low information entropy. Yeah. So the way these manifolds work the the LLM start paying attention to prompts that have high information entropy Yeah. and low prediction entropy. So what do I mean by that? So So when I say I'm going out for dinner Yeah. right? So when I say I'm going out for dinner that phrase the the LLMs have been trained you know, they've seen it a lot and there are many different directions I can go with it. I can say I'm going for dinner tonight, I'm going for dinner to McDonald's, or I'm going to dinner blah blah blah. There there are many different Yeah. But when I say I'm going to dinner with Martin Casado you know, the LLM now this is information-rich. This is sort of a rare phrase. And now the the sort of realm of possibilities reduces because Martin is only going to take me to Michelin-star restaurants. >> Yeah. Yeah. Uh I'm not going to go to a McDonald's. You know, you you you get what I'm saying. The moment you add more context Yeah. you make the prompt information-rich the prediction entropy reduces. Yep. Yep. Yep. Yep. And another example that uh I often Well, I mean, but but just quickly, what So what But what is your takeaway? What is your implication on that? Which is of course as as your intro So yeah, so you're um uh So sorry sorry, I forgot how you described it, but like so the the more precise you are, the more tokens you are, I presume the less options you have for the next token. Is that correct or not correct? Yeah. Yeah. Essentially. So you're redu you're reducing it you're reducing it to like a very specific state space when it comes to confidence in an answer. And like this is kind of a manifold that you can go on. And then I mean, do you do you have kind of a conclusion of what that means for systems or what that means for reasoning or is it just a nice way to articulate the bounds of LLMs? No, that that there is something uh I don't know I don't know if I should say profound, but but there is something about it which tells what these LLMs can or cannot do. Right? So it it uh one of the examples that uh I often tell is suppose I ask you what is 769 * 1,025? You have no idea. You can have some vague idea given the two numbers, right? And so in your mind the next token distribution of the answer is going to be diffuse. Right? You don't know. You have maybe a vague guess. If you are, you know, mathematically very good, maybe your guess is more precise, but it's still going to be diffuse and it's not going to be the correct answer. But if I if you say can I write it down and do it the way we have learned multiplication tables, now you know exactly what to do next step. Right? You write 769 and then 1025 and then you know exactly. So at each stage of that process your prediction entropy is very low. You know exactly what to do. Because you have been taught this algorithm. And by invoking this algorithm, saying, "Okay, I'm not going to just guess the answer, but I'm going to do it step by step. Then your prediction and entropy reduces. And you can arrive at an answer which you're confident of and which is correct. And the LLMs are pretty much the same way. You know, that's why chain of thought works. What happens with chain of thought is you ask the LLM to do something chain of thought, it starts breaking the problem into small steps. These steps it has seen in the past. It has been trained on. Maybe with some different numbers, but the concept it has been trained on. And once it breaks it down, then it's confident. Okay, now I need to do A, B, C, D and then I arrive at this answer. Whatever it is. Let's zoom back. I want to want to get into LLMs, but before first Vishal, maybe you can give more of context on your background and how that informs your your work here. Okay. So yeah, yeah, as my team said, my background is very similar to his. We, you know, we come from doing networking. So my PhD thesis, my sort of early work at Columbia has all been in networking. But there's another side of me, another hat that I wear. Which is both an entrepreneur and a cricket fan. >> I was going to say, don't you own a cricket team or something? I am a minority owner at your for your local cricket team, the San Francisco Unicorns. Yeah, that's right. Very proud to have you. So but the uh So so in the '90s I was one of the people who uh uh started this uh portal called Cricinfo. And uh Cricinfo uh at one point it was the most popular website in the world. It had more hits than Yahoo. That was before India came on. Remarkable. >> And so you know, uh we built Cricket is a very star-studded sport. You'll think baseball multiplied by a thousand. And we had built this free searchable stats database on cricket called Statsguru. And this has been available on on Cricinfo since 2000. But because you can search for anything, everything was made available on Statsguru. And you know, you can't expect people to write SQL queries to query everything. So how do you how did we do it? Well, it was a web form. You know, where you could formulate your query using that form that and in the back end that that was translated into SQL query, got the results, and got it back. But as a result, that because you could do everything, everything was made available, the web form had like 25 different check boxes, 15 text fields, 18 different drop downs. The interface was a mess. It was very daunting. So and ESPNcricinfo in the mid 2006, I think. But they still kept the same interface. And that has always sort of nagged me. And so I still know the people who run it. >> nagged What nagged you? Is that Cricinfo did not have a formal language, it had a web form for doing queries? That web form was terrible. [Laughter] Because of that only the real nerds used that form. >> in the world to bother you, the fact that an old website was a web form. It was I appreciate I appreciate your commitment to aesthetics. [Laughter] So so I I'm still friendly with the people who run ESPNcricinfo. They're the editor-in-chief. Whenever he comes to New York, you know, we meet up, we go out for a drink. And so he was here in 2000. So now the story shifts to how LLMs and me sort of met. So January 2000, right before the pandemic, he was here and I again said, "Why don't you do something about Statsguru?" And he looks at me and says, "Why don't you do something about Statsguru?" He was kind of joking, but uh he he thought maybe, you know, I had some ways to fix the interface. So anyway, then the pandemic hit, the world stopped. But in July of 2020, the first version of GPT-3 was released. And I saw someone uh use GPT-3 to write a SQL query for their own da- database using natural language. And I thought, "Can I use this to fix Statsguru?" So I got early access to GPT-3. You know, getting access those days was difficult, but somehow I got it. But soon I realized that, you know, no, I cannot really do it. Because Statsguru, the the back-end databases were so complex and if you remember GT GPT-3 had only a 2048 token context window. There was no way in hell I could fit fit the complexities of that database in that context window. And and GPT-3 also did not do instruction following at that time. But then in trying to solve this problem, I accidentally invented what's now called RAG. Where based on the natural language query, I created a database of natural language queries and structure sort of the structured queries. I created a DSL which then translated into a REST call to Statsguru. So based on the new query, I would look through my set of natural language queries. I had about 1500 examples and I would pick pick the six or seven most relevant ones. And then that and the structured query I would send as a prefix and the new query and GPT-3 magically completed it and the accuracy was very high. So that had been running in production since September 2021. You know, about 15 months before ChatGPT came. And you know, the whole revolution in some sense started and RAG became very popular. So I didn't call it RAG, but this is something sort of I accidentally did in trying to solve that problem for Cricinfo. Now once I once I built it, you know, I was thrilled that this worked, but I had no idea why it worked. You know, I stared at that I stared at that what transformer architecture diagram. I read those papers, but I couldn't understand how or why it worked. So then I started in this journey of developing a mathematical model trying to understand how it worked. So that that's been sort of my journey through this world of AI and LLMs because I was trying to solve this cricket problem. Yeah. Amazing. And and so maybe reflecting back since since the release of GPT-3, what has most surprised you about how LLMs have have developed? So what has most surprised me? The pace of development. So GPT-3 was, you know, it was a nice parlor trick and you had to jump through hoops to get it to do something useful. But starting with you know, ChatGPT was an advance over GPT-3 and then you had all these things like chain of thought, instruction following. GPT-4 really made it polished. And you know, the pace of development has really surprised me. Now, you know, when I started working with GPT-3, I could sort of see what its limitations were, what I could make it do, what I couldn't make it do. But I never thought of it as you know, what it what these LLMs have become for me now and what what I've become for millions of people around the world. We treat these models as our co-workers. Almost like an intern that, you know, you're constantly uh chatting with them, brainstorming, making them do all sorts of work which we couldn't imagine, you know, just when ChatGPT was released. You know, it was nice. It was it could write poem, it could write limericks, it could answer some hallucinated uh questions. But the capabilities that have emerged now, that pace has been very sort of surprising to me. Do you see progress plat- plateauing or how do you either now or or in the near future how how do you see it going? I yes, in some sense progress is plateauing. Uh it's like the iPhone, you know, when the iPhone came out, wow, what is this thing? And then and the early iterations, you know, constantly we were amazed by new capabilities. But the last, you know, seven, eight, nine years, it's maybe the camera got a little bit bit better or, you know, one thing changed here or memory is more, but there has been no fundamental advance in what it's capable of. You can sort of see a similar thing happening with these LLMs. And this is not true for just one one company and one model. Right? You look at what OpenAI is coming up with or what Anthropic, Google, or all these open source Chinese model or Mistral, the capabilities of LLMs has not fundamentally changed. They've become better, right? They've improved, but they have not crossed into a different realm. So this is something that I really appreciate about your work. And so um the thing that really struck me is as soon as these things showed up, you actually got busy trying to have a formal model of what they're capable of, which was in stark contrast to what everybody else was doing. Everybody else was like, "AGI, these things are going to, you know, recursively self-improve." Like or or or they'll say, "Oh, all are just stochastic parrots, which doesn't mean anything. So, everybody had rhetoric, and sometimes this rhetoric rhetoric was fanciful, and sometimes this rhetoric was almost reductionist, like, "Oh, it's just a database," which is clearly not true. And the thing that really struck me about your work is you're like, "No, let's figure out exactly what's going on. Let's come up with a formal model, and once we have a formal model, we can reason about what that means." And then, you know, in in my reading of your work, I kind of break it into pieces. There's the first one where you basically you came up with this, you know, matrix abstraction. I think it's worth you talking through. And then >> Yeah. you took in-context learning as an example, and you mapped it to Bayesian reasoning, which to me was incredibly powerful, cuz at the time, nobody knew why in-context learning worked. So, I think it'd be great for you to discuss that, because again, I think I think it was the first real kind of formal effect on like like, how are these things working? And then, the more recent work that you're working on now is a kind of more generalizing version of of of what is the state space that these models output when it comes to comes to confidence, which is the manifold that we're talking about uh before. So, I would it would be I think it'd be great if you just described your matrix model, and then how you use that to just to to to provide some bounds what in-context learning is doing. What's what's happening. Okay, so so so yeah, let's start with that matrix abstraction. So so the idea behind the matrix is you have this gigantic matrix where every row corresponds to a prompt. And then, the number of columns of this matrix is the vocabulary of the LLM, the number of tokens it has that it can emit. So, for every prompt this matrix contains the distribution over this vocabulary. Yep. So, when you say the cat sat on the you know, the column that corresponds to mat will have a high probability. Most of them will be zero. But, you know, reasonable ex- continuations will have a non-zero probability. And so, you can imagine that there's this gigantic matrix. Now, the size of this matrix is, you know, if you just take just the old uh first generation GPT-3 model, which had a context window of 2,000 tokens and a vocabulary of 50,000 next tokens or 50,000 tokens then, the size of it, the number of rows in this matrix is more than the number of atoms across all galaxies that we know of. So, clearly, we can't represent it exactly. Now fortunately a lot of these rows are do not appear in real life. Right? An arbitrary correct collection of tokens, you're not going to use that as a prompt. Similarly uh you so so a lot of these rows are absent, and a lot of the column values are also zero. Right? When you say the cat sat on the, it's unlikely to be followed by the token corresponding to, let's say, numbers. Or, you know, an arbitrary collection of tokens. There will be only a very small subset of tokens that can follow a particular prompt. So, this matrix is very very sparse. But, even after that sparsity, and even after removing the sort of gibberish prompts, the size of this matrix is too much for these models to represent, even with a trillion parameters. So, what in an abstract sense, what what is happening is the models get trained on certain, you know, data from the training set and certain some a subset a small subset of these rows you have reasonable values. For the next token distribution. Whenever you give the prompt something new like then, it'll try to interpolate with what it has learned and what's there in the new prompt, and come up with a new distribution. But, it's basically so it it's more than a stochastic parrot. It is sort of Bayesian on this uh uh on this subset of the matrix that it has been trained on. So so when I say, you know "I'm going out for dinner with Martin tonight." Now I'm reasonably sure that it has never encountered that phrase in its training data, right? But, it has encountered variants of this phrase. And given that I'm going out with Martin it it can produce a Bayesian posterior. It uses that evidence that Martin is the one that I'm going for dinner with, and it'll produce a next token distribution that'll focus on the likely places that we are going. So so this matrix because it's represented in a compressed way yet the models respond to everything, every prompt. How do they do it? Well, they they go back to what they've been trained on interpolate there, and use the prompt as sort of some evidence to compute a new distribution. Right. So right. So the the the context of the prompt impacts the posterior distribution. Ex- exactly. Yeah. Right. And this is and this is you you mapped to Bayesian learning where the the the context is the new evidence. New evidence, exactly. >> So so I I'll give you so so so for instance, uh the cricket example that I spoke about earlier. Yeah. So I created my own DSL. Yep. Which, you know, mapped a natural language query in cricket to this DSL which then I can translate into a SQL query or a REST API or whatever. But, getting the DSL is important. Now, these LLMs have never seen that DSL. I designed it. Yeah. Right? But, yet after showing a few examples, it learned it. How did it learn that? >> is this is in the prompt. You didn't no training the person. 100% in the prompt, right? So, like it's the way the waiter stand time. Yeah yeah, this this is this was happening in October of 2020. Right? I had no access to internals of OpenAI. I could just, you know, access the API. OpenAI had no access to internal structure of Stats Guru. Or the DSL that I cooked up in my head. Yet, after showing it only a few examples, it learned it right away. So, that's an example where it has seen DSLs or structures in the past. And now using this evidence that I show, "Okay, this is what my DSL looks like." Now, a new natural language query, it is able to create the right posterior distribution for the tokens that map to the example that I've seen. Now, the the other beautiful thing about this is this is an example of few-shot learning or in-context learning, right? But, when I give that prompt along with this these examples to this LLM I'm not saying to the LLM, "Okay, this is an example of few-shot learning, so learn from these examples." Right? You just pass this to the to the LLM as a prompt, and it processes it exactly the way it would process any other prompt, which is not an example of in-context learning. So, that really means that the underlying mechanism is the same. Right? Whether you give a set of examples, and then ask it to complete a task a task like in in-context learning, or just give it some prompt for continuation, that I'm going out for dinner with Martin tonight. There's no in-context learning there. But the the process with which it's generating or doing this inferencing is exactly the same. And that's what I have been trying to model and come up with a formal model of. What I've found very impressive is you've used this basic model to show a number of things, right? To describe in-context learning and to map it to Bayesian learning, but you did it for another one where you kind of you've sketched out this almost glib argument on Twitter, on X where you made this um uh you you made a rough argument for why recursive self-improvement can't happen without additional information. And so, maybe maybe just walk through very quickly how like the same model you can just very quickly show that a model can never self recursively self-improve. So, uh you know, another phrase that uh uh we've been using recently is, you know, the output of the LLM is the inductive closure of what it has been trained on. Yeah. So, when you say that it can recursively self-improve uh it could mean one of two things. So, let let's get back to the >> well, actually, you know what's kind of interesting is like often the I mo- most people agree that if you have one LLM, and you just feed the output into the input, like it's not going to do anything. But then, often people will say, "Well, what if you have two LLM you have no external information, but you have two LLMs talking to each other. Maybe they can improve each other, and then you can have like, you know, a takeoff scenario." But again, you even addressed this, even in the case of like n number of LLMs using kind of the matrix model to show that like you just aren't gaining any information. Uh Yeah, entropy, yeah. Yeah, so so so you can represent the the sort of information contained in these models. And let's go back to that matrix analogy that have the matrix abstraction. So like I said, you know, the these models are uh represent a subset of the rows. Right? Yeah. >> So a subset of the rows are uh represented. But some of these rows are able to help fill out some of the missing rows. For instance, you know, uh if the model knows how to do multiplication doing the step-by-step, then every row that is corresponding to let's say 769 * 125 or whatever, all those multiplications it can fill out the answer because it has those algorithms sort of embedded in them. You just need to unroll them. Yeah. So it can sort of self-improve up to a point. But beyond that point, uh these models can only uh sort of generate what they've been trained on. So let me give you I'll give you three examples. Yeah. So any model any LLM that was uh trained on pre-1915 physics would never have come up with a theory of relativity. Einstein had to sort of reject the Newtonian physics and come up with the space-time continuum. He completely rewrote the rules, right? So that is an example of, you know, AGI. Where you are generating or generating new knowledge. It's not simply about the universe, right? It's not computing the universe. It's actually discovering something fundamental about the universe. Fundamental. And for that you have to go outside your training set. Similarly, you know, any any LLM that was trained on it would not have come up with quantum mechanics. Right? That that's wave-particle duality or this whole probabilistic notion or that, you know, energy is not continuous but it is quantized. You had to reject Newtonian physics. Yeah. Or Gödel's incompleteness theorem. Yeah. He had to go outside the axioms to say that, "Okay, it is incomplete." So those are examples where you're creating new science or fundamentally new results. That kind of self-improvement is not possible with these architectures. They can refine these They can fill out these rows. Yeah. Where the answer already exists. Another example, you know, which has received a lot of press these days is these IMO results, International Math Olympiad. Yeah. You know, whether it's a human solving it or the LLM solving it, they're not inventing new kinds of math. Yeah. They are able to connect known results in a sequence of steps to come up with the answer. Yeah. So even the LLMs, what they're doing is they are exploring all sorts of solutions. In some of these solutions, they they start going on this path where their next token entropy is low. So that's where where I say they they are in that Bayesian manifold. >> Yep. Yep. Yep. Where you have this entropy collapse. And by doing those steps, you arrive at the at the answer. But you're not inventing new math. You're not inventing new axioms or new branch branches of mathematics. Yeah. You're sort of using what you've been trained on to arrive at that answer. Yeah. So those things LLMs can do, you know, they'll get better at it of connecting the known dots. Yeah. But creating new dots, I think we need an architectural advance. Yeah. So Martin was talking earlier about how the discourse, you know, was was either a stochastic parrot stochastic parrots or, you know, AGI or because it's all new. How are you How do you conceive of sort of the AGI discourse or or or even the the the concept? What does it mean to the extent that it's it's useful? How do you think about that? So so the way, you know, I think about it, the way we try to formulate in our papers is it's it's beyond the stochastic parrot, but it's not AGI. It's doing Bayesian reasoning over what it has been trained on. So it's it's it's a lot more sophisticated than just a stochastic parrot. How do you define AGI? Okay, so AGI, uh so how do I define AGI? So the way I would say that LLMs currently navigate through this known Bayesian manifold, AGI will create new manifolds. So right now these models navigate, they do not create. AGI will be when we are able to create new science, new results, new math. When an AGI comes up with a theory of relativity, I mean, it's it's it's an extremely high bar, but you get what I'm saying. It has to go beyond what it has been trained on to come up with uh new paradigms, new science, and that's that's my definition of AGI. Vishal, can you Do you think that based on the work you've done, can you bound the amount of data, computer, or data or compute that would be needed in order for it to to evolve? So so So what are the problems if if you just take LLMs as they exist? Is it There is so much data used to create them. To create a new manifold will need a lot more data just because of the basic mechanisms, right? Otherwise, it'll just kind of like, you know, get kind of consumed into the existing set of data. Like Have you found any bounds of of of what would be needed to actually evolve the manifold in a useful way or do you think we just need a new architecture? I personally think that we need a new architecture. The more data that we have, the more compute we have, we'll get maybe smoother manifolds. So it's like a map. Yeah, cuz cuz I mean there's there's there's this view that people have. They're like, "Well, Vishal, this is all this is all this is all, you know, good and well, but, you know, I could just take an LLM and I can give it eyes and I can give it ears and I can put it in the world and it'll gain information and based on that information, it'll improve itself. Um and therefore it can learn new things, but the counterpoint that I've always just intuitively thought to that is the amount of data used to train these things is so large. How much can you actually evolve that manifold given an incremental I mean, it's almost none at all, right? There has to be some other way to generate new manifolds that aren't evolving the existing one. I I completely agree. There has to be a new sort of architectural leap that is needed to go from the current, you know, just throwing more data and more compute. You know, it's going to plateau. It's it's, you know, the iPhone 15, 16, 17. And are are there any research directions that are promising in in your mind that might help us, you know, go beyond LLM limitations? But so so I mean uh again, I love LLMs. They are fantastic. And they are going to increase productivity like nobody's business. But I don't think they are the answer. So, you know, Yann LeCun famously says that uh LLMs are a distraction on the road to AGI. >> end. They're dead end to AGI. I don't think I'm not quite in that camp, but I I think we need a new new architecture to sit on top of LLMs to reach AGI. You know, a very basic thing, you know, what Martin just said, you know, you give them eyes and you give them ears, you make them multimodal, they of course they'll become more powerful. But you need a little bit more than that. You know, the the way human brains learns with with very few examples, that's not the way transformers learn. Yeah. Uh and you know, I'm not saying that we need to create an Einstein or a Gödel, but there has to be an architectural leap that is able to create these manifolds. And just throwing new data will not do it. It'll just smoothen out the already existing manifolds. Is that something So is is your goal to actually help like think through new architectures or are you primarily focused on putting formal bounds on existing architectures? A bit of both. I mean, the the former goal is the more ambitious one that uh everybody is chasing. And yeah, I I think about that constantly. Are are there any new even like uh sort of hints at a new architect or like have we started to make any progress on on on new architectures? Or is it Uh You you you know, um Yann has been pushing at this Jepa architecture. Yeah. Uh energy-based architectures, uh they they seem promising. The the way I have been sort of thinking about it is you know, you there's this uh set of a benchmark or the ARC prize. Yeah. Right? That Mark Mike Canoop and Francois Chollet have have And if you understand why the LLMs are failing on this test, maybe you can sort of reverse engineer a new architecture that'll help you uh succeed in that, right? Uh and I agree with a a lot of what several people say that, you know, language is great, but language is not the answer. You know, when I'm looking at catching a ball that is coming to me, I'm mentally doing that simulation in my head. I'm not translating it to language to figure out where it will land. I do that simulation in my head. So where you know one of the new architectures architectural things is how do we do how do we get these models to do approximate simulations? To test out that idea and whether to proceed or not. So So so yeah, we have you know another thing that I've always wondered about is did we develop as humans did we develop language because we were intelligent or because we developed language we accelerated our intelligence? So I I don't know which side of the camp you fall on that question. >> mean what's interesting is like you have these anecdotal examples of humans developing languages de novo that have been recorded, right? Like it's it's either what the watermelon or Nicaraguan sign language, right? Where there is these students that develop their own language without being taught and so that would suggest that language is follows intelligence. The problem is is they're all anecdotal, right? Like who knows if somebody didn't teach them sign language? Like nobody really knows there is no controls. So this is all these observational studies and there's so few of them you have to wonder if it's just kind of sloppy observation. And so I think that the question is still outstanding. Yeah. So I mean language definitely accelerated our intelligence. There's no question about that. Yeah. But which followed which we don't know. I view it as I view it as a I view it as a networking problem naturally which is once you have languages you can communicate. You can communicate you can store you can replicate, yeah. Yeah yeah. Exactly exactly, right. Cool. Um again this is kind of a wonky question but um Yeah. Uh you know what I think one thing that you've brought to the discourse and for those that are listening to this I really think that you should look up Vishal's work and read it. I just think it'll give you a really really especially if you have a systems background like a networking or systems background. It'll give you a really really good understanding of kind of the bounds on these. Um but like the toolkit that you draw from is like information theory and like more formal have you found that the AI community is receptive to this or is it like two different cultures two different planets trying to communicate and not a lot of common ground? Like how have you found like bringing like the networking view of the world to the AI realm? Some of them are receptive to it definitely. But you know uh these large conferences at the reviewing process it's so random and the kind of questions they ask you know I'm a modeling person. I like to model things. And you know I submitted one version of this work to one very famous uh machine learning or AI conference and the reviewer said okay this is a model so what? So So there is uh That's absolutely remarkable. So like you you've actually taken a system that nobody understands. We have no models for. You actually provided some model that we can use to analyze it and uh that alone wasn't sufficient. >> They ask and so where are the large scale experiments to to prove this? I do listen I I honestly I mean I I find there's so much empiricism in like the the the current you know AI community exactly cuz we don't understand the systems. You know it kind of reminds me I I I feel like I feel like systems went the other way, right? It's like we had all of these models but then we didn't understand how the systems worked and then we just like actually did measurement. It feels like ML and or the AI stuff is the opposite which is like we know we don't understand them and so we just measure them but now we're trying to like come up with the models. Yeah exactly. So it was so easy in some sense to build these uh artifacts and then just measure them that people have been going around trying to do that and you know one time I really dislike is prompt engineering. Why? >> You know engineering used to mean sending a man to the moon or providing five nines reliability. Prompt engineering is prompt twiddling. Yeah. You you fiddle with a prompt and the model changes and the the inference the output changes and you know you have like hundreds of papers just just you know doing one experiment the other changing a prompt this way that way and writing their observations. And as a result you know lots of these papers are being written are being submitted for review reviewers get busy looking at all this kind of empirical work. And my personal taste is to first try to understand model it. Yeah. And then you can do the other things. >> so like I'm a theory guy. I don't know about this bit twiddling like Let me ask one more LLM question which is Yeah. are there any benchmarks or real world tasks that if they if they occurred you'd sort of re-evaluate and say hey maybe LLMs are you know closer to the path to AGI than than I thought. If there were any real world tasks. Good question. You know which uh for LLMs uh all these models the one domain where you have the most training data is probably coding. And coding is where you can also have the most structure. And yet anyone who's used these tools whether it's cursor or whatever cloud code LLMs continue to hallucinate continue to generate unreasonable code. You know you have to you have to constantly uh babysit these models. So the day an LLM can create a large software project without any babysitting is the day I'll be a little bit convinced that it's towards AGI. But again I don't think uh it'll be able to create new science. If it does that's when I'll be convinced. I you know I think that you can almost take a definitional approach to answer this question Vishal like the problem with these types of questions is is if you have billions of dollars and you can collect whatever data you want you can make a model do anything you want, right? And so like you know what I'm saying like it's it's it's it's some level you've got this entire capital structure machinery behind these models. So you're like oh it can be good at science. Well sure you put a billion dollars to solving materials science and collect all this data you'll be good at material science or or whatever it is. And so but there is a definitional answer which is and and and and I'm going to draw from your work which is there is a manifold that's in there based on the data it's been training on and then the question is is if it ever produces something that's off like a new manifold so considering the existing training data if it ever does that. If it does something that's outside of that distribution then clearly we're on a path to to learning new things and if not then everything is just a computational step from what's already known. Yeah so so I mean And then I guess I guess the counter I guess the counter to that would be maybe all humans do is work on their own manifold and Einstein uh you know was lucky or something I guess would be the counter to that but I I just don't know. >> you know that's several many unseen examples and yeah it's creating this new manifold. I didn't want to use that definitional answer. I thought it might sound too Yeah. too wonky too mathematical. But essentially if LLMs really created this new manifold then I would be convinced. But so far they have just gotten better at navigating the existing manifold the existing training set. Which is hugely powerful and it's going to change the world. >> hugely I'm not denying that. I think they are extremely extremely good Yeah. at what they can do. But there's a limit to what they can do. So I have one quick question. What's next for you? I mean you've uh you you've you've tackled in context learning. You've got a model for LLMs and then you've got a generalized model for like you know like their solution space. What are you thinking about tackling next? Yeah in terms of uh modeling or Academically an LLM >> Academically yeah I academically I'm uh you know I'm I'm thinking of this what is the architectural leap that is needed Oh that's exciting. to create this new manifold and how do we use you know multimodal data? Awesome. to to expand the realm of science and talk to us. That's right. We'd love that. So I mean you know even with LLMs you know the in the paper we say that you can improve uh the inference by following this low or minimum entropy path. So so that's a very sort of small step that we are taking you know we are building and trading models that will do inference based on the entropic path. Yeah. By the way is is model probe still up? Token probe yeah yeah token probe is still up and and you can see actually the you know token probe is software that we built and thanks to Martin and A16Z's generosity it's running on your servers and anyone can go and test. And what we have done there is we actually show the entropy. Yeah. It is so inviting. I recommend anybody listening to this who's interested actually check out Token Prop. It literally shows you the Yeah, as you go along. It's it's remarkable. You know, so in context learning, you know, you create your new new DSL and you give it to the prompt and you can see the confidence rising with each new example, the entropy reducing. And that sort of is a validation of the model. You can see it sort of unfolding in right in front of your eyes. The Token Prop is here. All right, thanks. Thanks again. Uh Vishal, thanks so much for coming to the podcast. It's a great conversation. We appreciate it. Thanks for It was great fun. Thank you. Thank you so much again. [Music] [Music]

Original Description

LLMs have made tremendous progress in modeling human language. But can they go beyond that to make new discoveries and move the needle on novel scientific progress? We sat down with distinguished Columbia CS professor Vishal Misra to discuss this, plus why chain-of-thought reasoning works so well, and what real AGI would look like. Timecodes: 0:00 Intro 0:32 How LLMs and humans reason through manifolds 4:15 Token prediction, entropy & confidence 8:05 Chain-of-thought reasoning and entropy reduction 10:20 Vishal’s background 14:10 Inventing RAG 17:30 The rise of LLMs and the question of plateau 21:00 The Matrix Model / how prompts map to token distributions 28:10 Why LLMs can’t recursively self-improve 34:02 Defining AGI 38:25 Future architectures & multimodal intelligence 42:00 Modeling vs prompt engineering 47:20 What would prove AGI has arrived? 50:01 Closing thoughts Resources: Follow Dr. Misra on X: https://x.com/vishalmisra Follow Martin on X: https://x.com/martin_casado Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://x.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from a16z · a16z · 0 of 60

← Previous Next →
1 a16z Podcast | Money, Risk, and Software
a16z Podcast | Money, Risk, and Software
a16z
2 a16z Podcast | Wall Street's Most Hated Man -- A Conversation With Overstock.com's Patrick Byrne
a16z Podcast | Wall Street's Most Hated Man -- A Conversation With Overstock.com's Patrick Byrne
a16z
3 a16z Podcast | How Big Companies Can Get the Most From Silicon Valley
a16z Podcast | How Big Companies Can Get the Most From Silicon Valley
a16z
4 a16z Podcast | The Role of Academia in the Startup World
a16z Podcast | The Role of Academia in the Startup World
a16z
5 a16z Podcast | AMPLab, the Power of Open Source, and the Future of Systems Software
a16z Podcast | AMPLab, the Power of Open Source, and the Future of Systems Software
a16z
6 a16z Podcast | Dell + EMC -- Why the Python Just Ate the Cow
a16z Podcast | Dell + EMC -- Why the Python Just Ate the Cow
a16z
7 a16z Podcast | Belief -- An Interview with Oprah Winfrey
a16z Podcast | Belief -- An Interview with Oprah Winfrey
a16z
8 a16z Podcast | Holy Non Sequiturs, Batman: What Disruption Theory Is ... and Isn't
a16z Podcast | Holy Non Sequiturs, Batman: What Disruption Theory Is ... and Isn't
a16z
9 a16z Podcast | Boards and the Power of Networks
a16z Podcast | Boards and the Power of Networks
a16z
10 a16z Podcast | A Whirlwind Tour of Policy Issues in Tech
a16z Podcast | A Whirlwind Tour of Policy Issues in Tech
a16z
11 a16z Podcast | Beyond Lean Startups
a16z Podcast | Beyond Lean Startups
a16z
12 a16z Podcast | Blockchain vs/and Bitcoin
a16z Podcast | Blockchain vs/and Bitcoin
a16z
13 a16z Podcast | Quantum Leap
a16z Podcast | Quantum Leap
a16z
14 a16z Podcast | Artificial Intelligence and the 'Space of Possible Minds'
a16z Podcast | Artificial Intelligence and the 'Space of Possible Minds'
a16z
15 a16z Podcast | Fintech from the World's Financial Capital -- London
a16z Podcast | Fintech from the World's Financial Capital -- London
a16z
16 a16z Podcast | On Recent IPOs and Comparing Private vs. Public Valuations
a16z Podcast | On Recent IPOs and Comparing Private vs. Public Valuations
a16z
17 a16z Podcast | The Future of Food
a16z Podcast | The Future of Food
a16z
18 a16z Podcast | Data Down on the Farm
a16z Podcast | Data Down on the Farm
a16z
19 a16z Podcast | The Data Science of Food and Taste
a16z Podcast | The Data Science of Food and Taste
a16z
20 a16z Podcast | Using Social Tools to Build Homes for Those Most in Need
a16z Podcast | Using Social Tools to Build Homes for Those Most in Need
a16z
21 a16z Podcast | London Calling for Tech Done in a Different Way
a16z Podcast | London Calling for Tech Done in a Different Way
a16z
22 a16z Podcast | Building Tech Startups in a Place Where Tech Isn’t Everything
a16z Podcast | Building Tech Startups in a Place Where Tech Isn’t Everything
a16z
23 a16z Podcast | Nootropics and the Best Version of Your Brain, Yourself
a16z Podcast | Nootropics and the Best Version of Your Brain, Yourself
a16z
24 a16z Podcast | Scaling Ideas and Startups in the U.K. and Europe
a16z Podcast | Scaling Ideas and Startups in the U.K. and Europe
a16z
25 a16z Podcast | The Tiger and the Dragon -- On Tech and Startups in India and China
a16z Podcast | The Tiger and the Dragon -- On Tech and Startups in India and China
a16z
26 a16z Podcast | Telepresence and Tech for a Distributed Workforce
a16z Podcast | Telepresence and Tech for a Distributed Workforce
a16z
27 a16z Podcast | The Present State and Future Possibility of Virtual Reality
a16z Podcast | The Present State and Future Possibility of Virtual Reality
a16z
28 a16z Podcast | Writing a New Language of Storytelling with Virtual Reality
a16z Podcast | Writing a New Language of Storytelling with Virtual Reality
a16z
29 a16z Podcast | Mellody Hobson and Ben Horowitz Talk Investing, Career, and Star Wars!
a16z Podcast | Mellody Hobson and Ben Horowitz Talk Investing, Career, and Star Wars!
a16z
30 a16z Podcast | The Future of Software Development
a16z Podcast | The Future of Software Development
a16z
31 a16z Podcast | What Software Developers (and Therefore Every Company) Need
a16z Podcast | What Software Developers (and Therefore Every Company) Need
a16z
32 a16z Podcast | Making the Most of the Data That Matters
a16z Podcast | Making the Most of the Data That Matters
a16z
33 a16z Podcast | Harnessing the DevOps Movement -- Don’t Go Chasing Waterfalls
a16z Podcast | Harnessing the DevOps Movement -- Don’t Go Chasing Waterfalls
a16z
34 a16z Podcast | Nobody Discusses Work Software Outside of Work -- and Then There’s Slack
a16z Podcast | Nobody Discusses Work Software Outside of Work -- and Then There’s Slack
a16z
35 a16z Podcast | The Fundamentals of Security and the Story of Tanium’s Growth
a16z Podcast | The Fundamentals of Security and the Story of Tanium’s Growth
a16z
36 a16z Podcast | Things Come Together -- Truths about Tech in Africa
a16z Podcast | Things Come Together -- Truths about Tech in Africa
a16z
37 a16z Podcast | When Banking Works Like My Smartphone
a16z Podcast | When Banking Works Like My Smartphone
a16z
38 a16z Podcast | How to Be Original and Make Big Ideas Happen
a16z Podcast | How to Be Original and Make Big Ideas Happen
a16z
39 a16z Podcast | The Future of Money and Monetization
a16z Podcast | The Future of Money and Monetization
a16z
40 a16z Podcast | Building Affirm, and Why Max Levchin Has Watched Seven Samurai 100-Plus Times
a16z Podcast | Building Affirm, and Why Max Levchin Has Watched Seven Samurai 100-Plus Times
a16z
41 a16z Podcast | Hall of Fame Football Meets Venture Capital
a16z Podcast | Hall of Fame Football Meets Venture Capital
a16z
42 a16z Podcast | Breaking the Barriers of Human Potential
a16z Podcast | Breaking the Barriers of Human Potential
a16z
43 a16z Podcast | 'In the Eye of a Tornado': Views on Innovation from China
a16z Podcast | 'In the Eye of a Tornado': Views on Innovation from China
a16z
44 a16z Podcast | Infrastructure... Is Everything
a16z Podcast | Infrastructure... Is Everything
a16z
45 a16z Podcast | Mobile Falls Hard for Virtual Reality
a16z Podcast | Mobile Falls Hard for Virtual Reality
a16z
46 a16z Podcast | Disruption in Business... and Life
a16z Podcast | Disruption in Business... and Life
a16z
47 a16z Podcast | Data Network Effects
a16z Podcast | Data Network Effects
a16z
48 a16z Podcast | The Dream of AI Is Alive in Go
a16z Podcast | The Dream of AI Is Alive in Go
a16z
49 a16z Podcast | I Reject the Term Viral Video
a16z Podcast | I Reject the Term Viral Video
a16z
50 a16z Podcast | Truth and Humanity in Leadership
a16z Podcast | Truth and Humanity in Leadership
a16z
51 a16z Podcast | Your Worst Deeds Don’t Define You -- Life and Redemption in Prison
a16z Podcast | Your Worst Deeds Don’t Define You -- Life and Redemption in Prison
a16z
52 a16z Podcast | Investing in (Business and Career) Change
a16z Podcast | Investing in (Business and Career) Change
a16z
53 a16z Podcast | Scaling Companies and Culture
a16z Podcast | Scaling Companies and Culture
a16z
54 a16z Podcast | Teams, Trust, and Object Lessons
a16z Podcast | Teams, Trust, and Object Lessons
a16z
55 a16z Podcast | The Why, How, and When of Sales
a16z Podcast | The Why, How, and When of Sales
a16z
56 a16z Podcast | Selling to Developers & Open Source Business Models
a16z Podcast | Selling to Developers & Open Source Business Models
a16z
57 a16z Podcast | Connectivity and the Internet as Supply Chain
a16z Podcast | Connectivity and the Internet as Supply Chain
a16z
58 a16z Podcast | E-commerce, Payments, & More in India's Evolving Retail Landscape
a16z Podcast | E-commerce, Payments, & More in India's Evolving Retail Landscape
a16z
59 a16z Podcast | Banking on the Blockchain
a16z Podcast | Banking on the Blockchain
a16z
60 a16z Podcast | On Corporate Venturing & Setting Up 'Innovation Outposts'
a16z Podcast | On Corporate Venturing & Setting Up 'Innovation Outposts'
a16z

This video explores the potential of LLMs to achieve AGI, discussing their strengths and limitations, and the importance of chain-of-thought reasoning and prompt engineering in advancing LLM capabilities.

Key Takeaways
  1. Understand the basics of LLMs and their applications
  2. Learn about chain-of-thought reasoning and its role in LLMs
  3. Evaluate the limitations of LLMs in recursively self-improving
  4. Explore the concept of AGI and its requirements
  5. Consider the importance of multimodal intelligence in achieving AGI
💡 Chain-of-thought reasoning is a crucial component of LLMs, enabling them to make new discoveries and advance scientific progress, but LLMs are limited in their ability to recursively self-improve, highlighting the need for additional research and development to achieve AGI.

Related Reads

Chapters (14)

Intro
0:32 How LLMs and humans reason through manifolds
4:15 Token prediction, entropy & confidence
8:05 Chain-of-thought reasoning and entropy reduction
10:20 Vishal’s background
14:10 Inventing RAG
17:30 The rise of LLMs and the question of plateau
21:00 The Matrix Model / how prompts map to token distributions
28:10 Why LLMs can’t recursively self-improve
34:02 Defining AGI
38:25 Future architectures & multimodal intelligence
42:00 Modeling vs prompt engineering
47:20 What would prove AGI has arrived?
50:01 Closing thoughts
Up next
Solve Any Math Problem Step by Step — Free (Type or Snap a Photo)
Zariga Tongy
Watch →