AI, Robot
Skills:
Agent Foundations80%LLM Engineering70%Tool Use & Function Calling60%Multi-Agent Systems50%Autonomous Workflows50%
Key Takeaways
The video explores the world of artificial intelligence, covering topics such as AI agents, robotics, reinforcement learning, and embodied AI, with tools like DeepMind and OpenEye, and techniques like human feedback and simulation-to-reality transfer
Full Transcript
[Music] the idea of creating artificial creatures has obsessed us for millennia but for the purposes of this exercise before we go any further I'm going to ask you to purge any of the following thoughts from your head checklist Hephaestus expelled from Olympus and then built two servant robots the boot of ant a hunter or spirit movement machines of 12th century India made to protect the relics of Buddha the chess-playing Mechanical Turk just a bloke in a box pulling levers Mary Shelley's Frankenstein he walks he talks not a lot else kubrick's how don't call me Dave c-3po r2d2 canine ns2 and seriously these are just numbers and letters Robbie the robot robot lovers Robocop [Music] feeling better okay let's get going I'm Hanna Frey I'm an associate professor in mathematics and I am AI curious and this is deepmind the podcast series where we look at the fast moving story of artificial intelligence we've been talking to the scientists researchers and engineers based at deepmind in London and we're looking at how they're approaching the science of AI and some of the tricky decisions the whole field is wrestling with at the moment so whether you just want to know more or want to be inspired on your own AI journey then this is the place to be you see the thing is robots cell we've long lusted after the idea of appending the natural order with human ingenuity we just can't seem to leave it alone and in this episode we are looking at AI and robotics Marie Shanahan is a senior scientist at deep mind he's also a professor of cognitive robotics at Imperial College London and growing up Marie was utterly mesmerised by science fiction so you can picture his face when the Hollywood film director Alex Garland approached him following the publication of his book embodiment and the inner life Alex contacted me and said oh I'm writing a script for a film about AI and consciousness and I read your book and it you know helped to crystallize some ideas and you like to chat about about it and and so of course it was you know it was a great opportunity to get involved in a science-fiction film and then and then to my great good fortune it turned out to be an absolute cracker and that is how Murray became scientific advisor on the oscar-winning film ex machina I met with Murray to get a potted history of AI more people tend to think of AI is this very new thing a very a very modern invention but it's actually been around for quite a long time it has the idea of artificial intelligence the idea of making artificial creatures dates back to Greek mythology but there's a modern conception of AI perhaps really dates back to Alan Turing's paper published in the 1950s where he first kind of asked the question could a machine think and gave a number of kind of refutations for counter arguments to the idea that a machine could think this is where the Turing test this is that this famous paper inaugurated the so-called Turing test course cheering didn't call it the Turing test which is the idea that we should subject the machine to a test to see whether it's basically whether it's indistinguishable from a human in dialogue the term artificial intelligence was actually coined by John McAfee a Stanford professor he was at MIT at the time and John McAfee organized a conference in 1956 bringing together a lot of leading thinkers in maths to try and scope out the idea of building a thinking machine and he coined the term artificial intelligence what did they describe as at that time how did they see artificial intelligence John McCarthy in particular his idea of artificial intelligence that he had in mind was a kind of system that would answer questions really and be able to engage in dialogue with with humans so something that you're talking to although of course in those days it wouldn't have been through speech it would have been by typing in at the keyboard and it was very much a disembodied notion of artificial intelligence so this system didn't have a body and interact with a physical world in the way that we do or animals do or indeed that robots do so they weren't really thinking about robotics at that point I'm sort of imagining something like how in 2001 a Space Odyssey exact typing in Robbins yeah and under kind of nice version you know that their approach to artificial intelligence was to build systems that reasoned in logic and we now think of that whole sort of approach to artificial intelligence of using logic and reasoning as so-called good old-fashioned artificial intelligence or Oh classical AI now I know I know I did a bad thing there I mentioned how but nobody ever said that this was gonna be easy but to recap in classical AI you have to write down a complete list of rules for how you want your agent to think if this happens then do this if that happens then do that it's a nice idea in theory but if your agent is going to know how to handle every possible scenario you can throw at it it's going to need to be a long long list there was a project called psych which attempted to write out all the rules of common sense to build an enormous encyclopedic database of a common sense so I mean there'll be things like if you've got a container and you put something in that container and then move that container somewhere else then the thing that was in it gets moved as well then yeah stuff like that you know and if you buy something and pay for it you'll have less money than you had in the first place so do I think it's impossible to do that I think it's impossible in practice because it turns out that the sheer number of rules that you would have to write is absolutely enormous we might have moved away from these long lists of rules by now as a way to teach our artificial intelligence in favor of agents that can learn the rules for themselves but the skills that we want are AI to have like good old-fashioned common sense are just as important now as they ever were [Music] imagine one day long into the future a wealthy computer scientist builds an AI to manage his stamp collection he plugs into the internet gives it access to his bank account and sets in the challenge to buy as many stamps as possible at first the agent acts as its creator intended signing up to eBay bidding on stamps but after a while it gets another idea more money equals more stamps so why not start trading on the stock market to make more money and it soon realizes it can get the stamps cheaper if it can get them at source so the agent buys up a factory convert its manufacturing process to stamp making and goes on with achieving its goal but of course the limiting factor here is paper more paper more stamps so it starts commanding forests to be felled the wood to be processed all to feed its single-minded ambition more stamps now there's no denying that the AI is doing what it was told but it's doing so at any cost and any agent without some kind of common sense will be at risk of taking our instructions a bit too literally this might be a bit of an extreme example but Victoria crackovia a research scientist at deep mind working on AI safety is already seeing agents that aren't exactly behaving in the way their designers intended a reinforcement learning agent that was playing a boat racing game and the intended behavior there was to go around the racetrack and finish the race as soon as possible and the agent was encouraged to do this by having these little green squares along the track that would give its rewards and then with the agent figured out is that instead of actually playing the game it could be going around in circles and hitting the same green squares over and over again to rack up more points then you have this this whole situation with the boat going in circles and crashing into everything and catching and fire and still getting more points than would otherwise get these kind of situations are quite common because you haven't ever stopped the AI from doing that or told the AI that you don't want it to do that that's a perfectly sensible solution for it to come across yeah from the perspective of the AI it can't really tell that this solution is achieved it's just something that gives it a lot of reward so it can necessarily distinguish between the general solution and a really creative solution that humans just haven't foreseen there are plenty of examples like this one team of researchers created an agent inside a very simple two-dimensional computer game and tossed it with building itself a body to get itself from the start line to the finish line quite quickly it worked out that it could just build itself taller and taller and taller until it was as high as the track was long and then just flopped forwards to cross the line and there was the agent playing the game of Tetris which realized it could just pause the game forever and never lose but there is a balancing act here we don't want our AI misbehaving but we also don't want to restrict our agents too much and this is part of what's tricky about achieving safe behavior is that we don't want to hamper the system's ability to come up with really interesting and innovative solutions that we have not foreseen so we don't just want human imitation we want superhuman capabilities but without unsafe behavior there is a very fine line between a naughty algorithm and one that's finding innovative solutions to problems that humans haven't been able to solve the AI doesn't know the difference between the two it doesn't understand what's really important to us it doesn't have any common sense and that means you have to be very careful when you're setting up incentives and rewards for your agents part of the reason that this is so difficult is this general perfect that's called good hearts law and economics or good hearts law says is that when a metric becomes a target it ceases to be a good metric a classic example of God arts law comes from British India when authorities offered cash rewards for dead Cobras as a way to decrease the population of snakes unbeknownst to the British the local started to breed Cobras in order to take advantage of the reward as soon as the authorities found this out they scrapped the scheme altogether and revoked the rewards but now there were these snake farms everywhere filled with worthless Cobras so what did the locals do release the Cobras into the country resulting in an increase in the cobra population this is what the scientists call a specification problem when your specified objective fails to bring about the intended behavior this is generally quite likely to happen because you my preferences tend to be quite complex and whenever we try to distill them into some kind of specification or something that we say we want it's going to be a lot simpler than our real preferences are and wouldn't necessarily capture everything that's important to us let's imagine we're living in the future where robot Butler's are commonplace you clearly specify the objective for your robot it should serve you at all times now how is that agent going to feel about its own off switch your old always has an incentives to preserve its own function if it gets turned off they can vacuum the floors anymore can't bring you coffee it want to disable its off switch for example Yan Leica is a senior research scientist at deep mind also working on AI safety if you turn off then contact you the forest so if it understands how the opposite of switch mechanism works it would want to disable it what we want is you want our systems to actually do something that is good for us ready to do something that we actually wanted not just like me we want it but how do you get around these kinds of problems well we already know that writing a long list of do's and don'ts won't work however long your list gets you're always going to forget something the new breed of learning agents is going to need a different approach so one direction that we are pursuing is learning your what functions for reinforcement learning agents and you can thing kind of think of this as learning what your system should be doing from human feedback so for example in in one work that we did together with open the eye is where we trained a simulated robot to do a back flip and the way this works is like the the robot does some movement and then you watch a video of their movement and you say two videos and then you can compare which of those looks more like a back flip yawns experiment has a human sitting and watching a screen looking at an AI attempting to do a back flip each time the human will feedback on whether the attempt was good enough slowly nudging the agent in the right direction here is the key idea with constant human feedback the human can communicate their preferences without having to actually specify them and risk oversimplifying things in the process and after you like a few hundred rounds of feedback the robot can actually perform the black flip it's kind of learn what the objective should be that the objective should be your back flip and what a back flip is this is difficult to specify precisely what a back flip would be I mean so in my case like I can't do a back fin right and but I can see if the systems do back flips and in some ways like this allows us to get superhuman capability the AI then is essentially being rewarded by pleasing the human in a way exactly okay so we can we can do a little experiment I'll try to teach you to make a sound okay by giving feedback so you'll make two sounds mm-hmm and that you wish and Intel sounds is closer to you what I have in mind okay yeah does that sound good so I'm being a are you yeah you're being the I and I'm being the human teacher and you're training me essentially with reinforcement learning and my reward function is getting you to be happy you know what that's like so the reward function here is like something is in my mind right and I'm trying to teach you so you can't directly see the reward you can only see like my feedback but my objective is to get you to say you like the sound I'm making exactly okay listen think 80 cents let's go for meat and this went on for a while beep beep and but with Ian's feedback we got beep and beep I go with the first one absolutely nowhere the expiration problem actually if I was actually I have and I've gone through ten thousand iterations by now well you have a human the loop that you've use all of that so it is it is kind of slow but this is actually a problem that we have in our assistance is that like in order to give useful feedback you have to have useful examples right in this case you produced sounds that are very similar and like the son I had in mind was like very different so I really like have the opportunity to give you a very useful feedback but this is like basically this is the same thing would come up with a bad flip right like if your robot just like lies on the floor like two different ways like what are you gonna do but you would hope though after I'd had maybe a hundred goes at it I'd end up with something that began to approach what you had in mind yeah definitely there finally is it like oh no I'm not asking any questions I'm like I'm in AI I mean ideally this is at some point this is what we'd want our systems to be able to do I know just like describe this on to you in natural language and then you could just do it that would be really cool right yeah so this is the kind of research that we want to do in the future that's like the kind of system so we want to figure out how to build okay actually everything I've done so far has been like voices you didn't say that it had to be a vocal thing did you yeah okay hang on all right how about this how about a service and the first one the only problem was setting up reinforcement learning with a human standing over it offering feedback at every stage is that it is monumentally slow now it would take an agent months if not years to master a game like Atari and even then you need to hire a pretty large group of poor students to do the boring job of supplying feedback in real time but there is an alternative you can rustle up a slightly more sophisticated learning partnership so what we do when we actually build these systems is we don't literally do the experiment that we did now and instead we have we train an eel network a second meal network that learns how I would give feedback as the human and then the the neural network can teach you because I can just oversee all of the things you're doing by now neural networks have become really good at spotting patterns dog not dog backflip not backflip perhaps you don't always need a human in the loop laborious ly giving feedback why not have two agents one attempting the task and the other deciding if it succeeded the reason why this works very well is because evaluating the objective is an easier task than producing the behavior that achieves it so you can have like D for example and the backflipping robot all the neural network has to do is like look at what the robot is doing and see whether it's a back flip you're listening to deepmind the podcast a window on AI research but of course where this stuff really comes alive is where you take the ideas take them out of a computer simulation and allow your imagination to roam into the world of embodied AI robots that learn how to cook robots that learn how to pack fruit in crates tuck you into bed and perform backflips it's time to visit the deep mind robotics laboratory this is your lab this is where you spend your days this is jackie kay a software engineer packed to the rafters with robots it's very cool we're basically running out of the best it's quite noisy in here yeah we got a lot going on today this place isn't quite the high-security basement laboratory you'd imagine there are half assembled robot arms and machine parts scattered all around the place some look like mechanical hands others quite a lot like great kitchenaid style food mixers and curiously there is a SpongeBob SquarePants mascot hanging from the ceiling much of the work in this room focuses on getting robotic arms to learn how to do simple tasks each robot is bolted to the floor in its own cordoned off area and as we come in the door there's one that's caught my eye now that game that you play when you're a kid you have tennis bat and a ball attached to it and you kind of play tennis when you're wearing this sort of looks like a robotic version of that it's called the ball in a cup oh she is it is exactly that it is trying to swing the ball into that little basket so you can see it swinging the ball around the little it's actually tracking position of the yellow ball and it's trying to minimize the distance from that ball to the area inside the basket not exactly a smooth delivery of the ball into the cup more a slightly clumsy lucky flip of course if your ultimate goal was to make a perfect cup and ball playing robot you could directly program a machine to do that without fail you can build robots that perform simple tasks in much more elegant ways but that's not really the point here's why I had so senior research scientist a deep mind a robot is a machine that can take over a task one can say in the broadest sense that your dishwasher at home and your vacuum cleaner both robots because they do something complex they run on their own they have some autonomy to them of course we also like to think about robots as having some intelligence and then you get more into the realm of a robot with AI and I'm particularly interested in saying how can we take this a technology and make it work on robots so that we have embodied AI and that is an important distinction the focus of the research in this robotics lab isn't to create robots that are told what to do but to have them learn their own skills much in the same way that other agents do here the robots here are what's known as embodied AI so let's say yeah the task is a robot that's trying to lift a box is the programmer in a kind of traditional setting would say go to box open hand move hand you know 30 centimeters towards the box close hand but this is different yeah this is sort of taking what we want and then figuring out how to accomplish what we want every few seconds this particular robot will have an attempt at getting the ball into the cup before pausing resetting and having another go and every now and then by chance the ball lands in the cup and the robot is rewarded with a positive score just like if it was playing a computer game the reset in between their training episodes where it untangle itself or flips the ball out of the cup those are scripted but then when it actually tries to accomplish the task that is a policy which it has taught itself through experience over time from everything it learns and the schools it receives the robot builds a picture of how the ball moves and how this relates to the robots own movement so it had we come in right at the very beginning then what would we see just completely random noise probably very chaotic oh did he want ours it didn't wanna that was Wow that wasn't like three seconds so will it know now that that movement gave it a successful results yes and so the next time that we see this do we expect it to be better than when we walked in yeah it will try similar actions that gave it a positive reward or a success it's just so assessing that quite smartly right now yes there is a big advantage to getting the AI to figure out tasks for itself you'll end up with a robot that is much more flexible out of the box it doesn't matter what you want it to do tire not stacks and bricks peel a banana just as long as you can clearly communicate your objective you don't need to specify a long list of instructions for these robots and perhaps they'll come up with a new way of banana peeling that you haven't thought of but there's also a big drawback in training AI that has a physical body they are much slower to learn than all of the disembodied agents you'll find in this building the ones that only exist inside a computer other researchers are able to take advantage of parallelism that is they run their environments in simulation on computers they can run them in parallel on hundreds of compute different computers we're all gathering data about this environment they're trying to learn something about we only have what four robots here and there frequently not running the same experiment so it might just be one robot collecting data and that means our training will be orders of magnitude slower for comparable tasks compared to ones that are running in simulation on computers progress is slower in this room and you can tell the agents here are a little less accomplished in another corner another robot is trying to pick up a Lego brick with a hooked gripping device kind of like a claw and there's a rather ominous box of mangled Lego bricks next to it and I think this one has some shaping to close its gripper when it detects it's close to the brick it also has a fixed training time so after some number of seconds it will just get done go back to the start and then try it Lego is sort of this building block to general-purpose manipulation if you know if we can stack like two bricks together we can then do kind of I'm sure you're gonna break it oh it's fine sorry I got distracted by smashing up the legs there is real potential here and so Jackie and the team are constantly trying to find ways of speeding up that learning process one technique that we are looking into in order to and that some researchers here have done really cool work on in the past is something we call center wheel or simulation to reality transfer which shows are you where you take a simulation on a computer that models your robot and we can learn kind of in broad strokes what the robot is like how its actions affect its environment and how it can do something similar to a task that's trying to learn in real life so once we figure out all of that in simulation without even touching the real robot we can transfer the data it's collected onto real hardware so you can cheat basically you can cheat by imagining the real robot within a computer calculating all the physics that would happen in real life and then use the same techniques you use that army of computers to give you a bit of a head start before you even apply it to the real physical robot exactly so the real robots end up acting the same way as your simulated robots do well they start out acting the same way as their simulated robots do and then as they train more they might start behaving slightly differently or better when they go into reality using simulation might give you a head start on reality but it's never going to match precisely the real robot has to contend with grip friction gravity wear and tear all of which play important roles in the real world but none of which will be perfectly represented inside the computer and all of that means well I think science fiction may have set some false expectations one thing that's I am a little bit surprised but being here don't take this wrong way but these robots are a bit rubbish yes it's true I mean we've got a lot of work to do but as with so much of what happens here at deep mind it's not so much about these exact datums it's not about cup and ball or Lego stacking it's about the type of intelligence being acquired and how that fits into the bigger picture we want to demonstrate general purpose as we go along targets physical intentions great so contrasting that with kind of intelligence that's not embodied which maybe can learn leggins or maybe even understand language physical intelligence is looking at how physical actions of your body affect the real world so we want to take a wide variety of tasks playing with objects using tools maybe walking around in the future and we want to show that robots can teach themselves those how to do those tasks in the physical world in the physical world yes but physical intelligence is of course only one type of intelligence one string to the robots bow and deep mind as we've seen dares to dream big here's Marie Shanahan with the Big Finish the holy grail of AI research is to build artificial general intelligence so to build AI that is as good at doing an enormous variety of tasks as we humans are so so we are not specialists in that kind of way we you know a young adult human can learn to do a huge number of things you can learn to make food you can learn to make a company you can learn to build things to fix things you can do so many things to have conversations to rear children so all those things and we really want to be able to build AI that has the same level of generality as that if you want to know more about robotics and technical AI safety then head over to the show notes where you can also explore the world of AI research beyond deep mind and we'd welcome your feedback or your questions on any aspects of artificial intelligence that we're covering in this series so if you want to join in the discussion or point us to stories or resources that you think other listeners would find helpful then please let us know you can message us on Twitter or you can email us podcast at deepmind comm
Original Description
Selected as “New and Noteworthy” by Apple Podcasts, the highly-praised, award-nominated first season of "DeepMind: The Podcast" explores the fascinating world of artificial intelligence (AI). Join mathematician and broadcaster Hannah Fry as she meets world-class scientists and thinkers as they explain the foundations of AI, explore some of the challenges the field is wrestling with, and dives into the research that's led to breakthroughs like AlphaGo and AlphaFold. Whether you’re a beginner or an experienced researcher, join our journey into the past, present, and future of AI.
In this episode, Hannah explores the world of robotics. Do you ever find yourself confusing the science fiction of robots with the science of robotics? Looking beyond the sci-fi scenes with superintelligent robots that are uncannily human-like, the reality of building robots is more prosaic.
Inside DeepMind’s robotics laboratory, researchers are developing “embodied AI.” For example, robotic arms are learning from scratch how to pick up plastic bricks. But behind these simple actions are the cutting-edge challenges of bringing AI and robotics together. Listen as roboticists discuss training AI systems and robots to behave as intended, how we’re working to ensure our AI is safe, and why common sense is so important. 💡
“We want our systems to do something good for us, something we actually wanted, not just what we said we wanted.” – Jan Leike, Research Scientist at DeepMind
#robotics #feedback #safety #embodiedAI #physicalintelligence #DMpodcast
_ _
Listen to the full series on YouTube: http://dpmd.ai/3geDPmL
Or search for “DeepMind: The Podcast” on your favourite podcast app, including:
Search “DeepMind: The Podcast” and subscribe on your favourite podcast app.
Apple Podcasts: http://dpmd.ai/2Rzlmcu
Google Podcasts: http://dpmd.ai/3geDjp5
Spotify: http://dpmd.ai/3w29cb4
Pocket Casts: https://pca.st/30m1
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Google DeepMind · Google DeepMind · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
RL Course by David Silver - Lecture 8: Integrating Learning and Planning
Google DeepMind
RL Course by David Silver - Lecture 1: Introduction to Reinforcement Learning
Google DeepMind
RL Course by David Silver - Lecture 2: Markov Decision Process
Google DeepMind
RL Course by David Silver - Lecture 5: Model Free Control
Google DeepMind
RL Course by David Silver - Lecture 6: Value Function Approximation
Google DeepMind
RL Course by David Silver - Lecture 4: Model-Free Prediction
Google DeepMind
RL Course by David Silver - Lecture 3: Planning by Dynamic Programming
Google DeepMind
RL Course by David Silver - Lecture 10: Classic Games
Google DeepMind
RL Course by David Silver - Lecture 7: Policy Gradient Methods
Google DeepMind
Google DeepMind: Ground-breaking AlphaGo masters the game of Go
Google DeepMind
Match 1 - Google DeepMind Challenge Match: Lee Sedol vs AlphaGo
Google DeepMind
Match 2 - Google DeepMind Challenge Match: Lee Sedol vs AlphaGo
Google DeepMind
Match 1 15 min Summary - Google DeepMind Challenge Match
Google DeepMind
Match 3 - Google DeepMind Challenge Match: Lee Sedol vs AlphaGo
Google DeepMind
Match 2 15 Minute Summary - Google DeepMind Challenge Match 2016
Google DeepMind
Match 3 15 Minute Summary - Google DeepMind Challenge Match 2016
Google DeepMind
Match 4 - Google DeepMind Challenge Match: Lee Sedol vs AlphaGo
Google DeepMind
Match 4 15 Minute Summary - Google DeepMind Challenge Match 2016
Google DeepMind
Match 5 - Google DeepMind Challenge Match: Lee Sedol vs AlphaGo
Google DeepMind
Match 5 15 Minute Summary - Google DeepMind Challenge Match 2016
Google DeepMind
DQN SPACE INVADERS
Google DeepMind
DQN Breakout
Google DeepMind
Asynchronous Methods for Deep Reinforcement Learning: Labyrinth
Google DeepMind
Asynchronous Methods for Deep Reinforcement Learning: MuJoCo
Google DeepMind
Asynchronous Methods for Deep Reinforcement Learning: TORCS
Google DeepMind
Differentiable neural computer family tree inference task
Google DeepMind
StarCraft II DeepMind feature layer API
Google DeepMind
DeepMind Health – Partnership with the Royal Free London NHS Foundation Trust
Google DeepMind
DeepMind Health – Michael Wise – a patient's journey
Google DeepMind
Streams – a platform for a digital NHS
Google DeepMind
DeepMind Lab - Nav Maze Level 1
Google DeepMind
DeepMind Lab - Stairway to Melon Level
Google DeepMind
DeepMind Lab - Laser Tag Space Bounce Level (Hard)
Google DeepMind
Exploring the mysteries of Go with AlphaGo and China's top players
Google DeepMind
Demis Hassabis on AlphaGo: its legacy and the 'Future of Go Summit' in Wuzhen, China
Google DeepMind
The Future of Go Summit: AlphaGo & Ke Jie match 1 moves analysis
Google DeepMind
The Future of Go Summit: AlphaGo & Ke Jie match 2 moves analysis
Google DeepMind
The Future of Go Summit: Pair Go moves analysis
Google DeepMind
The Future of Go Summit: AlphaGo & Ke Jie match 3 moves analysis
Google DeepMind
Emergence of Locomotion Behaviours in Rich Environments
Google DeepMind
StarCraft II 'mini games' for AI research
Google DeepMind
Trained and untrained agents play StarCraft II full 1vs1 game
Google DeepMind
DeepMind open source PySC2 toolset for Starcraft II
Google DeepMind
ICML 2017: Test of Time Award (Sylvain Gelly & David Silver)
Google DeepMind
Ke Jie and DeepMind's Go Ambassador Fan Hui review the 3rd AlphaGo vs Ke Jie game
Google DeepMind
Ke Jie and DeepMind's Go Ambassador Fan Hui review the 1st AlphaGo vs Ke Jie game
Google DeepMind
Ke Jie and DeepMind's Go Ambassador Fan Hui review the 2nd AlphaGo vs Ke Jie game
Google DeepMind
AlphaGo Zero: Discovering new knowledge
Google DeepMind
AlphaGo Zero: Starting from scratch
Google DeepMind
Defining principles for tech companies in the NHS: DeepMind Health's Collaborative Listening Summit
Google DeepMind
A systems neuroscience approach to building AGI - Demis Hassabis, Singularity Summit 2010
Google DeepMind
Retour de Rémi Munos en France et ouverture de DeepMind Paris
Google DeepMind
Grid cells - Caswell Barry, UCL
Google DeepMind
DeepMind Health Research and Moorfields Eye Hospital NHS Foundation Trust: What our research shows
Google DeepMind
DeepMind Health Research and Moorfields Eye Hospital NHS Foundation Trust: A Patient's Story
Google DeepMind
Deep Learning 3: Neural Networks Foundations
Google DeepMind
Deep Learning 5: Optimization for Machine Learning
Google DeepMind
Deep Learning 8: Unsupervised learning and generative models
Google DeepMind
Reinforcement Learning 1: Introduction to Reinforcement Learning
Google DeepMind
Deep Learning 2: Introduction to TensorFlow
Google DeepMind
More on: Agent Foundations
View skill →Related Reads
📰
📰
📰
📰
Global Trade Dynamics Q3 2026 — Geopolitical & Macroeconomic Analysis
Dev.to AI
Bloomsbury among publishers in line for payout from Anthropic’s $1.5bn copyright settlement
The Next Web AI
Nvidia supplier Wistron opens $700m Texas factory to build Grace Blackwell superchips
The Next Web AI
University of Tennessee sues Anthropic over neural network patents
The Next Web AI
🎓
Tutor Explanation
DeepCamp AI