Accelerating Deep Learning with Mixed Precision Arithmetic, w/ Greg Diamos - #97

The TWIML AI Podcast with Sam Charrington · Advanced ·📐 ML Fundamentals ·8y ago

Key Takeaways

The video discusses accelerating deep learning with mixed precision arithmetic, featuring Greg Diamos, senior computer systems researcher at Baidu, who talks about the next generation of AI chips and how mixed precision arithmetic can improve performance in deep learning training. The conversation covers topics such as stochastic gradient descent, hierarchical representations, and the benefits of specialization in neural network architecture. The video also touches on the challenges of low preci

Full Transcript

[Music] hello and welcome to another episode of we'll talk the podcast where I interview interesting people doing interesting things in machine learning and artificial intelligence I'm your host Sam Cherrington last week I spent some time at CES the Consumer Electronics Show in Las Vegas exploring the vast sea of drones cameras paper-thin TVs robots laundry folding closets and other smart devices you name it it was there of course I was also able to sit down with some really interesting folks working on some pretty cool AI enabled products head on over to our YouTube channel to check out some behind-the-scenes footage from my interviews and other quick takes from the show and be on the lookout for our AI and consumer electronics series right here on the podcast coming soon the show you're about to hear is part of a series of shows recorded at the reworked deep learning summit in Montreal back in October this was a great event and in fact their next event the deep learning summit San Francisco is right around the corner on January 25th and 26th and will feature more leading researchers and technologists like the ones you'll hear on the show this week including Ian Goodfellow of Google brain and Daphne Koller of calico labs and more definitely check out the event and use the code twiddle AI for 20% off of registration in this show I speak with Greg diamond senior computer systems researcher at Baidu Greg joined me before his talk at the deep learning summit where he spoke on the next generation of AI chips Greg's talk focused on some of the work his team was involved in that accelerates deep learning training by using mixed 16-bit and 32-bit floating-point arithmetic we cover a ton of interesting ground in this conversation and if you're interested in systems-level thinking around scaling and accelerating deep learning you're really going to like this one of course if you like this one you're also going to liked when we'll talk 14 with Greg's former colleague kubo-san Gupta which covers a bunch of related topics if you haven't already listened to that one I encourage you to check it out and now on to the show hey everyone I am here at the rework deep learning conference in Montreal and I've got the pleasure to be seated across from Greg diamonds from Baidu and he's a senior researcher there Greg welcome to this week in machine learning in AI thanks for having me awesome why don't we get started by having you tell us a little bit about your background and how you got interested in ml nai sure absolutely my background is traditionally been in high-performance computing I've been really excited about building really fast processors for important applications that enable new applications that people can use you know it's actually kind of a strange story how I got to AI I used to be an AI skeptic okay that's always a good place to start yeah I always felt like AI would be valuable but I just felt like there's no way that simple algorithms like stochastic gradient descent could ever solve these complex you know highly multi-dimensional optimization problems and then at one point I remember sitting at an video research and hearing a talk from Yann LeCun and just realizing oh I was wrong I was totally wrong and immediately after that I joined by do research and rune was founding the Silicon Valley AI lab and it seemed like a great opportunity to learn more about AI and Mei had had spent such a crazy ride to get to this point I bet so what was your path to get to that point that you even had an opinion that involves to cast a gradient descent oh sure I mean well let's see I like building things that are useful for people uh-huh I feel like computing in general has enabled many you know new capabilities like the internet and like videos and you know so so many things that we take for granted every day but I really make our lives better and I was always really passionate about making computers even better okay it's kind of this belief that although you might not know it going into it if you make a faster computer or more efficient computer someone will find a way of building something amazing on top of that yeah and one of the and so I spend a lot of time looking at applications what are the things that you can use computers to do and AI was always there okay just the feeling with AI has you know for me before deep learning was that it would just be too hard that a lot of existing theory was kind of steering you down and and had all of these really difficult challenges and you know I'd seen a lot of people very smart people spend a lot of time trying to tackle those problems and not quite getting all the way there so do you recall what was it about young hawk that kind of made the light bulb go off and and made you realize that stochastic gradient descent was the the answer sure so you can look at these hierarchical future representations in confidence mm-hmm so when people you know look at images you can tell it's the world is hierarchical you can break a chair down into pieces it has arms and legs you can break those pieces down you know recursively and there's a lot of existing work that provides some evidence that vision algorithms will do similar things for recognition tasks that wasn't ever really a question and people were able to build by hand you know things like feature detectors and these hierarchical systems that worked reasonably well they were just very difficult to build and the interesting thing about yawns talk was that you know systems could do it automatically mm-hmm I thought before that that you would get stuck in these you know intractable optimization problems where even if a solution exists you know it's one out of some enormous ly large number like you know two to the the power of ten to the power of thirty you like something like you know you think about different sizes of numbers sometimes I think of the number of atoms in the universe as being a big number this is far far bigger than that and you were thinking that that number represents the represents what how many things you'd have to search through to find it the surface for your okay yeah it's like the needle in an enormous haystack more atoms than there are in the universe how could you ever possibly hope to search through efficiently Wow and especially with these very simple algorithms hmm but you know reliably I've seen since then for one application after another for image recognition for speech recognition for synthesis for language understanding to Brooks very reliably mm-hmm did you happen to catch any of Jeff intense talk glasses what do you think so is a kind of post SGD capsule well it's not post SGG actually the starting point is that SGD is really the only thing that we know works right but it was it's more posts kind of the traditional model of the neuron yeah I'd almost think of it as like post kava nuts okay and you know one thing we've realized recently is just there is a lot of complexity in modeling mm-hmm that while we like to think of deep learning as a general-purpose learning algorithm mm-hmm as you start applying it to different applications like my experience has been spending a lot of time applying it to speech recognition and you do get some benefits from more data and from some general-purpose aspects of the learning algorithm to the extent that it's robust to different speakers or different variations and different environments but you also get a lot of benefits from specialization so finding the right neural network architecture seems like it matters a lot and as we look into details for different applications as we spend time tuning neural network architectures for different applications you see very different structures emerge hmm so it wouldn't surprise me at all if there is a more efficient more general-purpose structure for vision than cough notes mm-hmm yeah one of the really helped me see that was a blog post by Steven merridy a while ago it may be almost a year ago at this point but he talked about network architecture being the new feature engineering Hindon stock was post calmness but he did start it off by talking about like calling into question the basic neuron structure but I don't he didn't necessarily he did kind of pivot to talking about the network architecture at a higher level right do you was there a piece in there where he suggested what might be the kind of successor to the traditional neuron architecture I think it's about this concept of capsules which might be groups of cooperating neurons okay and then the way that they cooperate together might be more more complex mmm let's see I feel like from a computational perspective it's actually hard to get away from the formulation of neurons that we have as the basic building block where even if there is something that's algorithmically more efficient or more well matched to the problem the computational building blocks that we have have been so highly tuned that if you made a very substantial change it might be better kind of at the element level right but it might be very inefficient on the type of computer that we know how to build today's breaks this whole ecosystem that we've built up around this traditional way of building neurons and networks and solving them yeah so much of the existing technologies are built on top of harder supported and also algorithm support for linear algebra Denson your algebra it's actually kind of surprising to me how effective that's been given that the primitives are so old and that they're so simple it's very surprising to me that those building blocks have gotten us as far as as far as they have so we took a little a little digression I guess before we got to kind of what you're up to EDI do sure so and Baidu really the Silicon Valley AI lab is a fouled building new breakthrough technologies and AI that especially have connections or enabled new products we focus on things that you know we can't currently do today and they're really multiple ways we end up attacking this the thing that I focus a lot on is just the idea of scale that as we have faster computers that can train larger more complex neural networks that it's not the only way but that's a very reliable way of improving accuracy or enabling new capabilities we've seen this in vision to a large extent it's actually kind of interesting when we started working in Baidu there is a question whether you could apply this outside of vision we spend a lot of time looking into speech recognition I think looking back on that it works very reliably like you can definitely apply dehorning outside of vision to many different applications sometimes now instead of you know thinking what is the new application that you can apply deep learning to I sometimes wonder are there any applications that are not well matched that we won't be able to make significant progress on by just applying this simple recipe of deep architectures large datasets large scale compute I haven't found one yet nice so we talked a little bit about before we got started we talked a little bit about the fact that you worked with one of our previous guests SIBO singh Gupta who was at Baidu and it's now at Facebook is that right yeah he's playing dota and he and I spoke pretty extensively about the speech translation the I forget the specific name of the project but the Baidu speech deep speech to deep speech right and in the course of that conversation we talked a little bit about you know some of the scalability challenges that your team ran into and tackling that problem but here at this conference you're talking about your you have a talk tomorrow in fact about some even further work that you've done can you tell us about what you're planning is out to talk about sure definitely so this is definitely along the lines of scaling deep neural networks we pretty consistently find that if you throw more data at the problem it isn't the only way but it is a very effective way of reducing error rates and improving accuracy and so often times when you keep throwing data at the problem you eventually run into some limitation sometimes you run out of data and sometimes you run out of patience to wait for your system to drain I've had a there's one example that I think drives drive this point home that we once had a model that bran I let it run on a large cluster run about 64 GPUs for about six months and now we were still getting improvements in accuracy at the time I decided to you know pull the plug now how if you're if you're still running your training model for oh you're saying the kind of your incremental error is decreasing as you ran it yeah this was a state-of-the-art model so it's actually improving the state of the art as it runs every minute is getting a little bit better than any model that we've had before and then after six months you really go back and look at that and say well I could let it run for another few years but I'd like to use it now so we're always looking for ways to improve speed there's actually you know oh man there's something that's related to that that I can't talk about yet we're a situation to sometimes be in where uh-huh like you know why something happens or you know that something will keep happening but you can't talk about it one of the things that I think I know is that for many applications we will continue to see improvements in the state of the art from faster computers I can't tell you why but I'm pretty sure I'll tell you why soon okay is this are we talking about a theoretical result or a okay yeah I think so you're nodding yes those who can't see I'm nodding yes yeah this is one of those things that'll definitely be surprising to people when we can finally talk about it but sorry I can't talk about today okay but you know you'll just have to take my word for it that we need faster computers okay and the talk tomorrow is gonna be about a way that we can make computers faster for deep warning okay yeah I spend a lot of my time thinking about this like what's the best that you could do how fast could you possibly make a computer even our understanding of physics on our existing technology I think one thing the industry is realizing is that we spend a lot of time focusing on general-purpose computation so building computers to run Windows or your browser but if you specialize if you build a computer that's created only a few things and not everything and you can do a lot better so we're exploring right now how do you build computers that are good at AI they're good at deboarding and is this different than what others in the space are doing like TP use and things of that nature let's see so this is about a very specific technique it might be one technology that might go into a chip like a GPA of these designs are using a lot of the same technologies okay this is a new one and this one has a pretty high upside this one has a maybe order of magnitude up side okay but before we dive into that can we take a second to kind of characterize the thing that GPUs and TP users are doing that's kind of gotten them the benefit and then we'll dive into you know this approach and what makes it different sure definitely so let's see I feel like one of the well they're two okay there are a lot of differences one thing that's worth keeping in mind is that modern processors incorporate probably thousands probably even more optimizations so these are technologies that will improve their performance in some way they might be circuit level they might be architecture level they might be in the software stack it's very hard because real designs are composed of you know you've pick out your favorite out of this pool of thousands of technologies and that becomes the new processor that you build these distinctions like TPU GPU or CPU they're very high level and they gloss over all of those details okay so I think what's more important than the name is what it does how fast is it actually like what is the result that you get from it how fast does it run a model that you care about mm-hmm we've seen things that were called CPU is being commonly used for training maybe 10 years ago there was a transition where people started using GPUs right the important thing about GPUs was the optimization for parallelism that there is abundant parallelism in neural network computations some of the things that are you know a few like a couple out of that list of thousand things that are being added into the next generation our optimizations around locality and low precision so the technology I'm going to talk about tomorrow is focused on low precision ok and there's a big difference when a lot of previous technologies have been discussed or proposed for low precision it's mostly been focused on inference and not training right and this will be one of the first results and especially the as far as I know largest scale result that focuses on using low precision for training ok I think the high level conclusion is it finally works it was enormous ly difficult Wow you know it was actually kind of a weird surprise that when you try doing low precision for training versus inference we didn't really know that we would see this we started looking into this but it just turns out that for some reason inference is so much easier than training that even you know very drastic reductions in precision like moving from double precision down to you know even 8-bit or possibly even lower fixed point representations it works just fine across many different models but if you try and do the same thing for training things fail it's actually kind of a funny point to me that we kind of made this implicit transition CPUs commonly support a high performance double precision GPU don't GPUs have historically optimized around single-precision so the 32-bit floating point instead of 64-bit floating-point it turns out it's kind of expensive to do this in a GPU to do 64-bit in a GPU where is this pretty achievement a cpu because you don't have to replicate this thing's this unit very many times right on a GPU have to replicate it a lot so if you replicate something big a lot it becomes expensive it's kind of surprising that the whole industry you know we when I started watching people train deep neural networks they might write scripts in you know MATLAB or you know call CPU libraries directly and those things by default used double precision when the industry switched to GPUs they switched from a double precision to single precision and we got so lucky it turns out that it didn't really matter but I think that was just by luck when we tried doing the next step we tried moving from single precision to half precision so moving from 32 bits and floating-point to 16 bit floating-point things started failing all over the place hmm and what caused those failures there were a lot of them let me try and draw a couple of big categories one was just differences in range so one of the points of having a floating-point as opposed to fixed point is that you have a very large dynamic range you might you know be able to do an operation like an add of a number that's you know where one number is a billion and the other number is you know ten to the minus five and that works right and so you need your range to extend from the smallest numbers that you want to deal with to the biggest numbers that you want to deal with and it turns out that if you look at all of the operations that go on in for propagation back propagation the nonlinearities in the SGD algorithm there's actually a pretty large dynamic range aren't we typically normalizing to try to get rid of some of that yeah it's interesting I'll come back to that plate up to that point yeah let me come back to that but I feel like the number one reason why when we just so the first experiments we did were just convert all of the 32-bit numbers to 16-bit precision numbers and try using exactly the same algorithm and also it by do you know because we were working on speech recognition we started doing this for recurrent neural networks mm-hmm it turns out that was one of the harder examples we started with one of the harder cases it turned out and so we would see all sorts of failures and what makes our n ends particularly harder I think it has to do with accumulated errors okay so as you keep doing this repeated application of a matrix multiplication with the same weights you're thinking about or overtime this just encourages extreme values either extremely small values or extremely large values this is sometimes people call this the vanishing gradients or exploding gradients aren't problems and for speech recognition we see very long time series okay we might see hundreds of iterations of an RN and or maybe thousands and we saw you know large accumulated errors over time mm-hmm one of the biggest sources of errors we came across was when you're actually combining gradients with the gradient update with the master copy of the weights so when you have this model it turns out it seems like SGD just makes these repeated small updates to a model and so if you look at it from a range perspective there's a large difference in magnitude between the magnitude of the gradients and the magnitude of the weights and so when you try and do operations on those you know numbers that have very different magnitudes you get loss of information or you get you get errors and that was one of the biggest problems we had with training and half precision it seemed like yeah moving for some reason the errors introduced from like the in floating-point things work out well if the numbers are in different magnitudes but not by too much and so it turns out that the difference for a single precision versus double precision was okay but it ended up being borderline for multiple applications when we were looking at the difference between single precision and half precision okay so we had to introduce some changes in order to deal with that one of the questions that came up in this previous conversation with shubo and which we tucked we touched on some of this stuff like I think pretty tangent it was like the end of our conversation I think if I remember correctly but we're talking about a reduced precision and I think I asked the question like you can reduce the the precision in multiple places you can reduce your the precision in your weights you can produce reduce the precision and there are outputs like when you're talking about reduced precision are you talking it sounds like you're talking about reduced precision everywhere just running on reduced precision infrastructure or in a reduced precision mode and not being particularly discriminating in terms of where you reduce the precision is that what you're referring to yes we're trying to keep it simple yeah we feel like if it ends up getting very complex and it's a it's difficult for people to know how you would actually apply this to a real model then you get back into your kind of architecture or feature engineering complexity issues yeah we definitely didn't want to introduce this as a number as another hyper parameter or right you know maybe this only works for a few layers but you know it doesn't work in these places and second you know you have to make this hard choice of deciding which ones to convert which ones not okay we wanted it just to be kind of like a switch and you would you know turn on the switch and you would get the performance improvement okay and I think we finally got to that point but for you know this this kind of reason there there were a lot of problems along the way so I mentioned the difference in magnitudes between the updates and the weights as a source of errors the other big one was just accumulated errors in long dot products so it turns out that taking weights convert quantizing them to 16-bit and then doing multiplications of activations with those weights didn't introduce too many errors but in neural networks especially in recurrent neural networks as layers get big you end up with these long dot products and so you're doing a running some kind of like over each row or for all of the inputs of a neuron and each operation has an accumulated error so every one in the sequence is gonna add some amount of error and now we're not talking about error in the kind of machine learning modeling since we're talking about floating point error yeah we're talking about just you know if you really wanted to do this multiply operation you didn't get the exact result we had to clamp it to a value that's representable by the grass shooter and so each time you do that you introduce quantization error and normally as long as you have enough bits the quantization error is small enough that it doesn't really affect the final result too much hmm exactly what too much means is very application dependent and into complex systems like neural networks it's really hard to know how much error is too much error other than just trying it on a real application right what we found for real applications like for speech recognition or for translation the error introduced by doing ads in 16-bit was too much error the models would diverge okay I more models would achieve significantly worse accuracy than the 32 bit baselines and so we went back to that and we tried a whole bunch of things like we tried hierarchical reductions and a bunch of things that ended up just being complicated and eventually we went back and looked at this circuits and came to the conclusion that it wasn't that expensive just to put in a 32-bit adder so you have a bunch of 16-bit multipliers and then you have a few 32-bit adders and if you look at the performance improvement that you get from that it ends up being most of the performance improvement that you would have got if you would have built 16-bit multipliers and 16-bit adders okay so yeah we ended with a mixed precision format you end up doing multiplication in 1316 pit but then the addition of 32-bit and there are a lot of other things we ended up looking at there's still some other failure cases but those are really the two big things as long as you keep the master copy of weights in 32-bit and as long as you do all the additions in 32-bit you can do all the multiplications and you can represent all the activations and intermediate copies of weights and weight gradients in 16-bit just to take a step back and make sure I understand why we're doing this are we talking about performance and computational costs are we talking about kind of unit compute costs for this chip by having you know narrower you know buses and things like that are we talking about training time performance like what are the the factors that are driving us to say we want to we want to do this and not just we can do it in reduced precision we want to do this in reduced precision sure yeah why do we want to do this in reduced precision it's really so we can build more efficient hardware with and without this technique you can just do a comparison if you're building the same processor with and without this technique there's a fair amount of performance at play it it might be something like four to eight X difference in really both sides of it total performance or energy per operation which would translate into efficiency okay so by going to reduce performance we can or by going to reduce precision we can increase you know some composite of performance and energy consumption by four to eight X like nearly order by order of magnitude yeah we could finish my six-month model and maybe just a single month mm-hmm and is it I guess I'm trying to get at this this question I don't know if the question makes sense but like is it the is it something inherent about the lower precision or is it the fact that the lower precision allows us to use new computer architectures that are faster in other ways or oh sure definitely so do you get this performance improvement on existing computer you get some performance improvement because you're moving around less data but it might be closer to 2x it really depends on whether your compute bound for bandwidth bound but the maximum might be more like 2x but if you build another computer okay if you build a new processor that was optimized around this idea you could do even better you can realize the 4 to 8 X okay so low precision fundamentally allows you to do an A train these neural nets by moving around less data right 16 bits instead of 64 for example so you get some advantage in doing that even if you're just in low precision mode on a general-purpose computer but it also allows you to build chips that are specific to running and low precision and that gives you that's where you get the big opportunity to bump up your speeds yes exactly okay and so you were here talking about the actual chip is that correct oh yeah so we're gonna talk about the voltage GPU from Nvidia this would say yeah this was a collaboration with Nvidia it's worth noting you know this hardware has been shipping for a while right but the side of it that we're talking about now is the validation that we've done on it so we've had always shown that you can train models in low precision we've looked at you know over 15 large-scale complete and and deporting applications so it's really easy to build hardware that gets great performance numbers but isn't able to run any real algorithms so from the point of view of loped of low precision the volta is like it's general-purpose right it's not a chip that's specifically designed for low precision oh it did actually have they're called tensor cores what was the name for them is a tensor core that is this operation I'm talking about okay it's a specialized unit that does 16 bit multiplication and floating-point with 32 bit floating point addition that unit wasn't designed as a result of this study got it got it and now if I remember correctly when this was announced they made a big deal about not the floating-point side of things but like in eighth formants and things like that how does that all fit in sure definitely so I kind of alluded to this maybe in the beginning that inference just is easier for some reason that okay training so that's all the in front that's like we can do in eight on in front side and it is easy and it just works and it's faster exactly got it okay I don't know I I don't know that this whole topic has been really fully explored yet uh-huh maybe someday in the future we might see someone who gets in date training to work but as far as I know I've never seen it I know there are a lot of there's a lot of work on you know very reduced precision like even down to binary but one thing is worth noting about these approaches is that they either have accuracy losses so you trade precision for accuracy on the complete application or they only apply to inference and not to training mm-hmm okay got it so reduce precision you did some validation that shows that essentially running in this mode is kind of a generalized approach you can take now you know things that you need to do or switches that you need to flip when you're training your model in order to get it to work accurately or to work correctly sure so one switch that you need to flip is you need to decide to do this okay you need to decide to represent things in sixteen bit and do your matrix multiplications or convolution operations in this mixed 16-bit 32-bit format that's somewhat of a global switch you can just turn that on for the entire program okay the other thing that you need to do that we found is essential is you need a master copy of the weights so in your optimization algorithm like your implementation of SGD you need to have a separate copy in 32-bit of all the weights and only when you're doing forward propagation or back propagation do you convert from that into 16-bit okay but both of those changes we found so the first one's really easy the second one can be encapsulated inside of the optimization algorithm so at least when you're designing a network you don't have to think about this okay you know I can envision how I might do this if I was writing the you know if I was implementing SGD myself to the higher level frameworks and toolkits all know how to do this or is that you know yet forthcoming so it's straightforward to do this in most frameworks but there needs to be developers who are working on those frameworks who will actually add support for this okay you know when we did this in the framework that we have invite you it's something like a you know 15 lines of code change so it's really minor but you still have to do that otherwise you won't get access to the improved performance right okay anything else you talked about in your or I keep saying it in past tense anything else you're going to talk about in your talk that you want to mention I guess the last thing is that this is one piece this is one technology that gives us a large improvement in performance I think we're aware of a lot of them that haven't been realized yet mm-hmm so you know I mentioned before the hardware industry for a long time has been kind of creeping along it's actually very difficult to realize large improvements in sequential performance at least for parallel programs like graphics applications and things that would run on GPUs performance has been increasing you know following something like the popular form of Moore's Law so exponential growth for AI if you're only thinking about running deep neural networks you can probably do a lot better than that in the short term so we might not have to wait 10 years to get a thousand X faster it might happen in just a few years I'm gonna mention some of the other ways that haven't been implemented yet but that we know about and are likely to happen in the future can you rattle those off this podcast will not be published before your talk tomorrow sure one of them is array parallelism I think this is one of the other big ones as array parallelism so okay I had a Forbes article about this for I was talking about locality the importance of locality if you build processors around the idea of locality and not just parallelism you end up with something that it looks I call it like an array processor rather than a vector processor okay you're thinking about the core instruction that you're doing instead of adding or multiplying two vectors together you're adding or multiplying two arrays together mm-hmm so you see things like this in designs like the TPU I feel like the thing that is wrong with those designs is that they don't find the knee of the curve that this is beneficial but you don't have to go all in on it to get most of the benefits and you actually are trading off so you don't find the knee of the curve what exactly does that mean it means working on arrays is a good idea but they don't have to be enormous arrays okay and you actually are trading off flexibility for performance when you're making the arrays and bigger so you shouldn't make them enormous you should make them big enough to get most of the savings and energy in terms of order of magnitude are we talking about like little teeny ones like convolutional kernel sizes are we talking about you know something bigger than that or yeah it's like more like 16 by 16 than 256 by 256 cents okay yeah yeah interesting any others on that list that come to mind one of the ones that doesn't work yet but I think is very promising is sparsity working with sparse representations or other than dense representations and it might seem like that's incompatible with the one that I just said the array parallelism mm-hmm we haven't shown this yet but I suspect that they're not incompatible so what I'm hearing putting the two together is that you know we're living in a world that thinks about all this stuff as composite vector operations and by thinking about this at the level of matrices you know there are opportunities there yes yeah that's a good way of thinking about it interesting well I really enjoyed this chat thank you so much for taking the time glad to be here okay thanks great all right everyone that's our show for today thanks so much for listening and for your continued feedback and support thanks to you this podcast finished the year as a top 40 technology podcast on Apple podcasts my producer says that one of his goals this year is to crack the top ten and to do that we will need your help please head on over to the podcast app rate the show hopefully we've earned your 5 stars leave us a glowing review and share it with your friends family coworkers Starbucks barista Zuber drivers everyone every review and rating goes a long way so thanks in advance for more information on Greg or any of the topics covered in this episode head on over to Twilio comm slash talks last 97 of course we'd be delighted to hear from you either via a comment on the show notes page or via twitter at at Twilio thanks once again for listening and catch you next time [Music] you

Original Description

In this show I speak with Greg Diamos, senior computer systems researcher at Baidu. Greg joined me before his talk at the Deep Learning Summit, where he spoke on “The Next Generation of AI Chips.” Greg’s talk focused on some work his team was involved in that accelerates deep learning training by using mixed 16-bit and 32-bit floating point arithmetic. We cover a ton of interesting ground in this conversation, and if you’re interested in systems level thinking around scaling and accelerating deep learning, you’re really going to like this one. And of course, if you like this one, you’re also going to like TWiML Talk #14 with Greg’s former colleague, Shubho Sengupta, which covers a bunch of related topics. This show is part of a series of shows recorded at the RE•WORK Deep Learning Summit in Montreal back in October. This was a great event and, in fact, their next event, the Deep Learning Summit San Francisco is right around the corner on January 25th and 26th, and will feature more leading researchers and technologists like the ones you’ll hear here on the show this week, including Ian Goodfellow of Google Brain, Daphne Koller of Calico Labs, and more! Definitely check it out and use the code TWIMLAI for 20% off of registration. The notes for this show can be found at twimlai.com/talk/97.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from The TWIML AI Podcast with Sam Charrington · The TWIML AI Podcast with Sam Charrington · 0 of 60

← Previous Next →
1 Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
The TWIML AI Podcast with Sam Charrington
2 How to Build Confidence as an ML Developer with Siraj Raval - #2
How to Build Confidence as an ML Developer with Siraj Raval - #2
The TWIML AI Podcast with Sam Charrington
3 Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
The TWIML AI Podcast with Sam Charrington
4 Interactive AI, Plus Improving ML Education with Charles Isbell - #4
Interactive AI, Plus Improving ML Education with Charles Isbell - #4
The TWIML AI Podcast with Sam Charrington
5 Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
The TWIML AI Podcast with Sam Charrington
6 Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
The TWIML AI Podcast with Sam Charrington
7 Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
The TWIML AI Podcast with Sam Charrington
8 Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
The TWIML AI Podcast with Sam Charrington
9 Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
The TWIML AI Podcast with Sam Charrington
10 Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
The TWIML AI Podcast with Sam Charrington
11 Building AI Products with Hilary Mason - #11
Building AI Products with Hilary Mason - #11
The TWIML AI Podcast with Sam Charrington
12 Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
The TWIML AI Podcast with Sam Charrington
13 Understanding Deep Neural Networks with Dr. James McCaffery - #13
Understanding Deep Neural Networks with Dr. James McCaffery - #13
The TWIML AI Podcast with Sam Charrington
14 Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
The TWIML AI Podcast with Sam Charrington
15 Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
The TWIML AI Podcast with Sam Charrington
16 Machine Learning in Cybersecurity with Evan Wright - #16
Machine Learning in Cybersecurity with Evan Wright - #16
The TWIML AI Podcast with Sam Charrington
17 Interactive Machine Learning Systems with Alekh Agarwal - #17
Interactive Machine Learning Systems with Alekh Agarwal - #17
The TWIML AI Podcast with Sam Charrington
18 Location-Based Intelligence for Smarter Marketing with Klustera - #18
Location-Based Intelligence for Smarter Marketing with Klustera - #18
The TWIML AI Podcast with Sam Charrington
19 AI-Powered Customer Support with HelloVera - #18
AI-Powered Customer Support with HelloVera - #18
The TWIML AI Podcast with Sam Charrington
20 Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
The TWIML AI Podcast with Sam Charrington
21 Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
The TWIML AI Podcast with Sam Charrington
22 Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
The TWIML AI Podcast with Sam Charrington
23 From Particle Physics to Audio AI with Scott Stephenson - #19
From Particle Physics to Audio AI with Scott Stephenson - #19
The TWIML AI Podcast with Sam Charrington
24 Selling AI to the Enterprise with Kathryn Hume - #20
Selling AI to the Enterprise with Kathryn Hume - #20
The TWIML AI Podcast with Sam Charrington
25 Engineering the Future of AI with Ruchir Puri - #21
Engineering the Future of AI with Ruchir Puri - #21
The TWIML AI Podcast with Sam Charrington
26 Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
The TWIML AI Podcast with Sam Charrington
27 Introducing Psycholinguistics into AI with Dominique Simmons- #23
Introducing Psycholinguistics into AI with Dominique Simmons- #23
The TWIML AI Podcast with Sam Charrington
28 Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
The TWIML AI Podcast with Sam Charrington
29 Offensive vs Defensive Data Science with Deep Varma - #25
Offensive vs Defensive Data Science with Deep Varma - #25
The TWIML AI Podcast with Sam Charrington
30 Global AI Trends with Ben Lorica - #26
Global AI Trends with Ben Lorica - #26
The TWIML AI Podcast with Sam Charrington
31 Intelligent Autonomous Robots with Ilia Baranov - #27
Intelligent Autonomous Robots with Ilia Baranov - #27
The TWIML AI Podcast with Sam Charrington
32 Reinforcement Learning Deep Dive with Pieter Abbeel  - #28
Reinforcement Learning Deep Dive with Pieter Abbeel - #28
The TWIML AI Podcast with Sam Charrington
33 Robotic Perception and Control with Chelsea Finn  - #29
Robotic Perception and Control with Chelsea Finn - #29
The TWIML AI Podcast with Sam Charrington
34 Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
The TWIML AI Podcast with Sam Charrington
35 The Power of Probabilistic Programming with Ben Vigoda - #33
The Power of Probabilistic Programming with Ben Vigoda - #33
The TWIML AI Podcast with Sam Charrington
36 Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
The TWIML AI Podcast with Sam Charrington
37 Video Object Detection at Scale with Reza Zadeh - #34
Video Object Detection at Scale with Reza Zadeh - #34
The TWIML AI Podcast with Sam Charrington
38 Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
The TWIML AI Podcast with Sam Charrington
39 Expressive AI-Generated Music With Google's Performance RNN with Doug Eck  - #32
Expressive AI-Generated Music With Google's Performance RNN with Doug Eck - #32
The TWIML AI Podcast with Sam Charrington
40 Smart Buildings & IoT with Yodit Stanton - #36
Smart Buildings & IoT with Yodit Stanton - #36
The TWIML AI Podcast with Sam Charrington
41 Deep Robotic Learning with Sergey Levine - #37
Deep Robotic Learning with Sergey Levine - #37
The TWIML AI Podcast with Sam Charrington
42 Deep Learning for Warehouse Operations with Calvin Seward - #38
Deep Learning for Warehouse Operations with Calvin Seward - #38
The TWIML AI Podcast with Sam Charrington
43 Cognitive Biases in Data Science with Drew Conway - #39
Cognitive Biases in Data Science with Drew Conway - #39
The TWIML AI Podcast with Sam Charrington
44 Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
The TWIML AI Podcast with Sam Charrington
45 Web Scale Engineering for Machine Learning with Sharath Rao - #40
Web Scale Engineering for Machine Learning with Sharath Rao - #40
The TWIML AI Podcast with Sam Charrington
46 Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
The TWIML AI Podcast with Sam Charrington
47 Machine Teaching for Better Machine Learning with Mark Hammond - #43
Machine Teaching for Better Machine Learning with Mark Hammond - #43
The TWIML AI Podcast with Sam Charrington
48 LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber  - #44
LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber - #44
The TWIML AI Podcast with Sam Charrington
49 Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
50 Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
The TWIML AI Podcast with Sam Charrington
51 Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
The TWIML AI Podcast with Sam Charrington
52 Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online  Meetup
Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
53 Word2Vec & Friends with Bruno Gonçalves -#48
Word2Vec & Friends with Bruno Gonçalves -#48
The TWIML AI Podcast with Sam Charrington
54 Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan  - #49
Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan - #49
The TWIML AI Podcast with Sam Charrington
55 Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
The TWIML AI Podcast with Sam Charrington
56 Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
The TWIML AI Podcast with Sam Charrington
57 AI-Powered Conversational Interfaces with Paul Tepper - #52
AI-Powered Conversational Interfaces with Paul Tepper - #52
The TWIML AI Podcast with Sam Charrington
58 Topological Data Analysis with Gunnar Carlsson - #53
Topological Data Analysis with Gunnar Carlsson - #53
The TWIML AI Podcast with Sam Charrington
59 ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
The TWIML AI Podcast with Sam Charrington
60 Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
The TWIML AI Podcast with Sam Charrington

The video teaches how to accelerate deep learning with mixed precision arithmetic, discussing the benefits and challenges of this approach, and how it can be applied to improve performance in deep learning training. The conversation covers various topics related to deep learning, including stochastic gradient descent, hierarchical representations, and the importance of locality in processor design. By watching this video, viewers can learn how to apply mixed precision arithmetic to their own dee

Key Takeaways
  1. Understand the basics of mixed precision arithmetic
  2. Apply mixed precision arithmetic to deep learning models
  3. Optimize deep learning training with specialized computation
  4. Use Nvidia's Volta GPU with tensor cores for mixed precision arithmetic
  5. Implement mixed precision arithmetic in most frameworks with minor code changes
  6. Flip a switch to represent things in 16-bit format
  7. Convert weights to 32-bit format for optimization
💡 Mixed precision arithmetic can improve performance in deep learning training by a large margin, but it requires careful consideration of the trade-offs between precision and performance.

Related Reads

📰
Best AI Classes in Indore | Join Today
Join AI classes in Indore to gain practical experience and industry-ready skills
Medium · Machine Learning
📰
Top 10 Machine Learning & Deep Learning Companies Transforming Enterprises in 2026
Discover the top 10 machine learning and deep learning companies driving enterprise innovation in 2026
Medium · Machine Learning
📰
Top 10 Machine Learning & Deep Learning Companies Transforming Enterprises in 2026
Discover the top 10 machine learning and deep learning companies driving enterprise innovation in 2026
Medium · Deep Learning
📰
Modal — Deep Dive
Learn how Modal Labs provides a high-performance, Python-native compute platform for data scientists, ML engineers, and AI researchers, and how it can help eliminate infrastructure tax
Dev.to AI
Up next
AI & Machine Learning Course Review by Tandeep Sandhu, Solutions Directior
Great Learning
Watch →