deeplearning.ai's Heroes of Deep Learning: Andrej Karpathy
Skills:
Fine-tuning LLMs90%LLM Foundations80%Supervised Learning80%Unsupervised Learning70%ML Maths Basics60%
Key Takeaways
The video discusses Andrej Karpathy's journey in deep learning, from his first exposure to restricted Boltzmann machines to his work on ImageNet classification, and explores the concepts of fine-tuning, supervised learning, and unsupervised learning, with tools such as TensorFlow, CommNets, JavaScript, and Python being mentioned
Full Transcript
so welcome Andre I'm really glad you could join me today yeah thank you for having so a lot of people already know your work in deep learning but not everyone knows your personal story so like to Austin start by you know telling us how did you end up doing all this work in deep learning yeah absolutely so I think my first exposure to deep learning was when I was an undergraduate at the University of Toronto and so Jeff Fenton was there and he was teaching a class on be performing and at the time it was a restricted Boltzmann machines trained on amnesty Jets and I just really like the way I'm kind of Jeff I talked about training the network like the mind of the network and he was using these terms and I just thought it would there was a kind of a flavor of something magical happening when when this was training on those digits and and so that's kind of like my first exposure to it although I didn't get into it in a lot of detail at that time and then when I was doing my master's degree at University of British Columbia I took a class with meta decorators and that was again on machine learning and that's the first time I kind of dealt kind of deeper into these networks and so on and kind of what was interesting is that I was very interested in artificial intelligence and sweat two classes and artificial intelligence but a lot of what I was seeing there was just very not satisfying like it was a lot of kind of you know death or surge breadth-first search alpha beta pruning and all these things and I was like not understanding out like I was not satisfied and so when I was seeing neural networks for the first time like and in machine learning which is kind of this term that I think is more technical and not as well known in kind of you know most people talk about artificial intelligence machine learning was more kind of a technical term I would almost say and so I was dissatisfied with artificial intelligence and when I saw machine learning I was like this is the AI that I want to kind of spend time on this is what's really interesting and and that's kind of what took me down those directions is that this is kind of almost a new computing paradigm I would say because normally humans write code but here in this case we write the optimization rights code and so you're creating input-output specification and then you have lots of examples of it and then the optimization rights code and sometimes it can write code better than you and so I thought that was just a very new kind of way of thinking about programming and that's what kind of intrigued me about it then through a work one of the things you've come to be known for is that you are now the human benchmark for the image net emission classification competition how did that come about so basically their image net challenge is kind of a it's sometimes compared to kind of the world cup of computer vision so a lot of people kind of care about this benchmark and number our error rate kind of goes down over time and it was not obvious to me in kind of where a human would be on the scale and I've done a similar smaller scale experiment on see part 10 dataset earlier so what I did then see part 10 is I just was looking at these 32 by 32 images and I was trying to classify them myself at the time this was only 10 categories so it's fairly simple to create an interface for it and I think I had an error rate of about 6 percent on that and so that was and then based on what I was seeing and how hard task was I think I predicted that the lowest error rate we've achieved would be like okay I can't exactly merge I think I guess like 10% and we're now down to like 3 or 2% or something crazy so that was my first kind of a fun experiment of like human baseline and and I thought it was really important for the same purposes that you kind of point out in some of your lectures I mean you really want that number to understand you know how well humans are doing in so we could compare machine learning algorithms to it and for imagenet it seemed that there was a discrepancy between how important is mention was and how much focus there wasn't getting a lower number and us not understanding even how humans are doing on this benchmark and so I created this a JavaScript interface and I was showing myself the images and then the problem with image net is you don't have just ten categories you have a thousand and so it was call most like a UI challenge of obviously I can't remember a thousand categories so how do I make it so that it's something fair and so I listed out all the categories and I gave myself examples of them so for each image I was crawling through a thousand categories and just trying to kind of you know see based on the examples I was seeing for each category what this image might be and I thought it was a just an extremely instructed exercise by itself I mean I was not I did not understand that like a third of omission at his dogs and like dog species and so that was kind of interesting to see that the network spends a huge amount of time caring about dogs I think a third of its performance comes from dogs and yeah so this was kind of something that I did for maybe a week or two I put everything else on hold I thought it was kind of a very fun exercise I got a number in the end and then I thought that one person is not enough I wanted to have multiple people and so I was trying to organize within a lab to get other people to kind of do the same thing and I think people are not as willing to contribute say like a week or two of like pretty painstaking work you know just like yeah sitting down for like five hours and trying to figure out which dog treat this is and so I was not able to get like enough data in that respect but we got at least like some approximate performance which I thought was was fun and then this was kind of picked up and it's uh it wasn't obvious to me at the time I just want to know the number but this became like a thing and people really liked the fact that that this happened and I'm referred to jokingly as like the reference human and of course I that's kind of a hilarious to me yeah well you what were you surprised when you know software defense finally surpass your performance absolutely so yeah absolutely um I mean especially I mean sometimes it's really hard to see in the image where it is is just like a tiny blob of like a black black dog is obviously somewhere there and I'm not seeing like you know I'm guessing between like 20 categories and the network just gets it and I don't understand how that comes about so there's some super humanists to it but also for the I think the network is extremely good at these kind of like statistics of like four types and textures and just I think in that respect I was not surprised that the network could better measure those fine statistics across lots of images in many cases I was surprised because some of the images required you to read like it's just a bottle and you can't see what it is but it actually tells you what it is in text and so as a human I can read it and it's fine but the network would have to learn to read to identify the object because it wasn't obvious just from from it um you know one of the things you've become well-known for and that the deep learning committee has been grateful to you for has been your teaching the Santa Clause and putting now online tell me a bit about how that came about yeah absolutely so I think I felt very strongly that basically this technology was transformative and that a lot of people want to use it it's almost like a hammer and what I wanted to do I was in a position to randomly kind of hand out this hammer to a lot of people and I just found that very compelling it's not like necessarily advisable from the perspective of a PhD student because you're putting research on hold I mean this became like hundred and twenty percent of my time and I had to put all of research on hold for maybe I mean I taught the class twice and each time it's maybe four months and so that time is basically spent entirely on the class so it's not super advisable from that perspective but it was basically the highlight of my PhD is not even likely related to research I think teaching the class was definitely the highlight of my PhD just just seeing the students just the fact that they were really excited it was a very different class normally you're being taught things that were discovered in 1800 or something like that but we were able to come to class and say look there's this paper from like a week ago or even like yesterday and there's new results and I think the undergraduate students and any other students they just really enjoyed that aspect of the class and the fact that they actually understood so there's not you know so you don't have to this is not nuclear physics or rocket science this is like you need to know calculus and linear algebra and you can actually kind of understand everything that happens under the hood and so I think just the fact that it's so powerful the fact that it's that it keeps changing on a daily basis if people kind of felt like they're on the forefront of something big and I think that's why people like really enjoyed that class a lot yeah and and and you've really helped a lot of people and a lot of hammers yeah you know as someone there's been and doing deep learning for quite some time now the field is evolving rapidly a bit crazy here how is your own thinking how is your understanding of deep learning change over these you know many years yeah it's basically like when I was seeing we're circling Boltzmann machines for the first time on digits it wasn't obvious to me how this technology was going to be used and how big of a deal it would be and also when I was starting to work on computer vision convolutional networks they were around but they were not something that a lot of the computer vision community kind of anticipated using anytime soon the I think the perception was that this works for small cases but would never scale to large images and that was just extremely incorrect and so basically I'm just surprised by how general technology is and how good the results are that was my largest surprise I would say and it's not only that so that's one thing that it worked so well and say like imagenet but the other thing that I think no one saw coming or at least for sure I did not see coming is that you can take these pre train networks and that you can transfer you can fine-tune them on arbitrary other tasks because now you're not just solving imagenet and you need millions of examples this also happens to be very general feature extractor and I think that's kind of a second inside that I think fewer people saw coming and you know there were there were these papers that are just like here all the things that people have been working on in computer vision I seen classification and action recognition object recognition you know face attributes and so on and people are just kind of crushing each task just by fine-tuning the network and so that to me was very surprising yeah and somehow I guess supervised learning gets most of the press and even though retraining fine tuning transfer learning is actually working very well people seem to talk less about the episode right yeah yeah I think what has not worked as much as some of these hopes are on unsupervised learning which I think has kind of been really why a lot of researchers have gotten into the field in around 2000 in 2007 and so on and I think the promise of that has still not been delivered and I think I I found that I find that also surprising is that the supervised learning part worked so well and the answer prize learning is still kind of in a state of uh-huh yeah it's still not obvious how it's going to be used or how that's going to work even though a lot of people are still deep believers I would say to use the term in the in this area so I know that you know one of the presidents wasn't thinking a lot about the long-term future of AI share your thoughts on that so I spent the last maybe year in a half at opening I kind of thinking a lot about these topics and it seems to me like the field will kind of split into two trajectories one will be kind of a kind of applied AI which is kind of just making these neural networks training them mostly with supervised learning potentially unsurprised learning and getting better say image recognizer zones or something like that and I think the other will be kind of artificial general intelligence directions which is kind of how do you get neural networks that are entire kind of dynamical system that thinks and speaks and can do everything that a human can do and as intelligent in that way and I think that what's been interesting is that for example in computer vision the way we approached it in the beginning I think was wrong in that we tried to break it down by different parts so we were like okay humans recognize people humans recognize scenes you can recognize objects so we're just going to do everything that humans do and then once we have all those things and now we have like different areas and once we have all those things we're gonna figure out how to put them together and I think that was kind of a wrong approach and we've seen that how that kind of played out historically and so I think there's something similar that's going on slightly on the higher level of with AI so kind of people are asking well okay people plan people do experiments to figure out how the world works or people talk to other people so when you language and people are trying to decompose it by function accomplish each piece and then put it together into some kind of brain and I just think it's kind of a just incorrect approach and so what I've been a much bigger fan of is having not decomposing that way but having a single kind of neural network there is the complete dynamical system that you're always working with a full agent and then the question is how do you actually create objectives such that when you optimize over the weights to make up that brain you get intelligent behavior out and so that's kind of been something that I've been thinking about a lot at opening I I think there are a lot of kind of different area ways that people have thought about approaching this problem so for example going in a supervised learning direction I have this essay online it's not an essay it's kind of a short story that I wrote and the short story kind of tries to come up with a hypothetical world of what it might look like if the way we approach this AGI is just by scaling up supervised learning which we know works and and so that gets into something that looks like Amazon Mechanical Turk or people association to lots of robot bodies and they perform tasks and then we train on that as a supervised learning data set to imitate humans and what that might look like and so on and so then there are other directions like on surprised learning from algorithmic information theory things like a ixi or from artificial life things that look more like artificial evolution and that's kind of where I spend my my time thinking a lot about and I think I had a correct answer but I'm not willing to reveal it here so you've already given out a lot of hammers and today there are a lot of people still wanting to enter the field of AR entity learning so for people in that position what advice do you have for them yeah so I think when people talk to me about CS 231 and and why they thought it was a very useful course what people what I keep hearing again and again is just people appreciate the fact that we got all the way to the low-level details and they were not working with a library they saw the route code and they saw how everything was implemented and implemented chunks of it themselves and so just going all the way down to the and understanding everything under you and never it's really important to not abstract away things like you need to have a full understanding of the whole stack and that's where I learned the most myself as well when I was learning this stuff it's just implementing it myself from scratch was the most important it was the piece that that I felt gave me the best kind of a bang for the buck in terms of understanding so I wrote my own library it's called comm Nijs it was written in JavaScript and implements convolutional neural networks that was my way of learning about back propagation and so that's something that I keep advising people is that that you not work with tensorflow or something else you can work with it once you have written it something something yourself on the lowest detail you understand everything under you and now you are comfortable to you know it's possible to use some of these frameworks that abstract some of it away from you but you know what's under the hood and so that's been something that helped me the most that's something that people appreciate the most when they take 231 and and that's when I would devise a lot of people yeah yeah and it's some kind of a sequence of layers and I know that when I add some drop out layers it makes it work better like that's not what you want in that case you're you're not going to be able to debug effectively you're not going to be able to improve on models effectively you know what dads are really glad that you've learned on the iCore starts off with that many weeks of Python programming first yeah good good thank you very much for sharing your insights and advice you're already a hero to many people in the deep learning world so really glad really grateful you could join us here today yeah thank you for having
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from DeepLearningAI · DeepLearningAI · 7 of 60
1
2
3
4
5
6
▶
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Forward and Backward Propagation (C1W4L06)
DeepLearningAI
deeplearning.ai's Heroes of Deep Learning: Yuanqing Lin
DeepLearningAI
deeplearning.ai's Heroes of Deep Learning: Ruslan Salakhutdinov
DeepLearningAI
deeplearning.ai's Heroes of Deep Learning: Yoshua Bengio
DeepLearningAI
deeplearning.ai's Heroes of Deep Learning: Pieter Abbeel
DeepLearningAI
deeplearning.ai's Heroes of Deep Learning: Ian Goodfellow
DeepLearningAI
deeplearning.ai's Heroes of Deep Learning: Andrej Karpathy
DeepLearningAI
Using an Appropriate Scale (C2W3L02)
DeepLearningAI
Gradient Checking (C2W1L13)
DeepLearningAI
Gradient Checking Implementation Notes (C2W1L14)
DeepLearningAI
Learning Rate Decay (C2W2L09)
DeepLearningAI
Understanding Mini-Batch Gradient Dexcent (C2W2L02)
DeepLearningAI
Mini Batch Gradient Descent (C2W2L01)
DeepLearningAI
The Problem of Local Optima (C2W3L10)
DeepLearningAI
Exponentially Weighted Averages (C2W2L03)
DeepLearningAI
Tuning Process (C2W3L01)
DeepLearningAI
Understanding Exponentially Weighted Averages (C2W2L04)
DeepLearningAI
Bias Correction of Exponentially Weighted Averages (C2W2L05)
DeepLearningAI
Gradient Descent With Momentum (C2W2L06)
DeepLearningAI
Normalizing Activations in a Network (C2W3L04)
DeepLearningAI
Hyperparameter Tuning in Practice (C2W3L03)
DeepLearningAI
Adam Optimization Algorithm (C2W2L08)
DeepLearningAI
RMSProp (C2W2L07)
DeepLearningAI
Fitting Batch Norm Into Neural Networks (C2W3L05)
DeepLearningAI
Why Does Batch Norm Work? (C2W3L06)
DeepLearningAI
Batch Norm At Test Time (C2W3L07)
DeepLearningAI
Softmax Regression (C2W3L08)
DeepLearningAI
Deep Learning Frameworks (C2W3L10)
DeepLearningAI
Neural Network Overview (C1W3L01)
DeepLearningAI
Training Softmax Classifier (C2W3L09)
DeepLearningAI
Why Deep Representations? (C1W4L04)
DeepLearningAI
Gradient Descent For Neural Networks (C1W3L09)
DeepLearningAI
Neural Network Representations (C1W3L02)
DeepLearningAI
TensorFlow (C2W3L11)
DeepLearningAI
Activation Functions (C1W3L06)
DeepLearningAI
Explanation For Vectorized Implementation (C1W3L05)
DeepLearningAI
Getting Matrix Dimensions Right (C1W4L03)
DeepLearningAI
Understanding Dropout (C2W1L07)
DeepLearningAI
Building Blocks of a Deep Neural Network (C1W4L05)
DeepLearningAI
Why Non-linear Activation Functions (C1W3L07)
DeepLearningAI
Computing Neural Network Output (C1W3L03)
DeepLearningAI
Backpropagation Intuition (C1W3L10)
DeepLearningAI
Train/Dev/Test Sets (C2W1L01)
DeepLearningAI
Deep L-Layer Neural Network (C1W4L01)
DeepLearningAI
Random Initialization (C1W3L11)
DeepLearningAI
Other Regularization Methods (C2W1L08)
DeepLearningAI
Normalizing Inputs (C2W1L09)
DeepLearningAI
Derivatives Of Activation Functions (C1W3L08)
DeepLearningAI
Parameters vs Hyperparameters (C1W4L07)
DeepLearningAI
Vectorizing Across Multiple Examples (C1W3L04)
DeepLearningAI
What does this have to do with the brain? (C1W4L08)
DeepLearningAI
Dropout Regularization (C2W1L06)
DeepLearningAI
Vanishing/Exploding Gradients (C2W1L10)
DeepLearningAI
Basic Recipe for Machine Learning (C2W1L03)
DeepLearningAI
Bias/Variance (C2W1L02)
DeepLearningAI
Forward Propagation in a Deep Network (C1W4L02)
DeepLearningAI
Weight Initialization in a Deep Network (C2W1L11)
DeepLearningAI
Numerical Approximations of Gradients (C2W1L12)
DeepLearningAI
Regularization (C2W1L04)
DeepLearningAI
Why Regularization Reduces Overfitting (C2W1L05)
DeepLearningAI
More on: Fine-tuning LLMs
View skill →Related Reads
📰
📰
📰
📰
The Bitter Lesson, Sweet Results: How We Rebuilt Picnic’s Recipe Recommender
Medium · Machine Learning
Training-Serving Skew: The Silent Bug That Kills Production ML Models
Medium · Machine Learning
The Magic Behind PyTorch: Why Autograd Is the Superpower Every Deep Learning Engineer Must…
Medium · Deep Learning
Machine Learning Pipelines Explained: Automating Your Entire ML Workflow
Medium · Machine Learning
🎓
Tutor Explanation
DeepCamp AI