Probabilistic Numeric CNNs with Roberto Bondesan - #482

The TWIML AI Podcast with Sam Charrington · Advanced ·🧬 Deep Learning ·5y ago

Key Takeaways

The TWIML AI Podcast discusses Probabilistic Numeric Convolutional Neural Networks with Roberto Bondesan, exploring Gaussian processes and probabilistic descriptions of discretization error, as well as other research topics like Adaptive Neural Compression and Gauge Equivariant Mesh CNNs.

Full Transcript

[Music] all right everyone i am here with roberto bandezan roberto is an ai researcher at qualcomm roberto welcome to the twimla ai podcast hi sam thank you i'm really looking forward to our chat we'll be talking about your paper at iclr probabilistic numerical numeric cnns uh but before we dig into that i'd love to hear you share a little bit about your background and how you came to work in ai sure so my background is in physics and that's how i got into ai by basically applying deep learning to some physics problems so before joining qualcomm i was working on characterizing new phases of matter and their potential for quantum computation so phase of matter uh is basically one of the states that in which you you can find you know matter or materials around you and one classical example is provided by the easy model so in the easy model you have you know some binary units some bits which can be either up or down and you can find this uh this system in two phases one order phase which is uh you know at low temperature where all these uh units we call them spins point in the same direction in a high temperature disorder phase where you these units point in random directions it turns out that characterizing you know phases of material is a very challenging problem and so researchers in the field they started to use ai for that so that's how i got into ai and as i was you know just learning about the techniques that were being developed this was around 2017-2018 for these problems i got very interested in ai per se and then i decided a career shift and to move indeed to do research in ai at this point it was also when the paper by taco coin and maxwelling who are my colleagues at qualcomm their paper on spherical cnns came out and so at that point i understood that you know studying neural networks using insights from physics was a very useful and interesting thing and that's what what i wanted to to work on so did you get were you familiar with the spherical cnn paper before you came to work with taco and max yes yes definitely that's indeed one of the links that brought me to to qualcomm yeah oh that's fantastic uh so tell us a little bit about your research interest yeah so after i after i joined qualcomm i i got involved into quite a few exciting projects so um one of them has to deal with the neural uh data compression so here the problem is that you want to compress data to you know send them over uh to some receiver and you can actually use neural networks in particular some forms of auto encoder to do that so that's a quite you know interesting problem from the theoretical point of view it involves generative modeling and things like that but it is also super impactful because you know it can immediately translate into uh different ways to to compress data so that's a very interesting uh direction otherwise i'm also very excited about quantum ai so after joining qualcomm i also had the the chance to develop some ideas around quantum deep learning so this is the you know intersection between quantum computation and ndi and so here is where also my my background in quantum thesis was quite useful to you know get started uh quickly on this field which is a rapidly developing field which i think is one of the most interesting you know directions which can disrupt ai in the future and finally uh more recently i got interested in also applying machine learning to combinatorial optimization problems so you know in industry we have a lot of combinatory optimization problems to solve and so the potential of using machine learning to you know improve on the classical techniques is very interesting and again this is a very interesting you know area between uh you know very uh theoretical uh mathematical work which you know is a is interesting to me and also very impactful work so yeah and in summary i had the the chance you know the of uh both doing uh impactful work for your society or say at large but also you know being able to um do some long-term uh research like quantum ai so i'm very excited about the you know a few things that are going on at qualcomm right now nice and when you say combinatorial optimization problems like traveling salesmen and map coloring and that kind of thing like classical combinatorial problems yeah precisely so in fact you know in the last few years uh maybe three four years um there's been a quite you know good progress on using deep learning for these problems so you know uh traveling assessment indeed this is the standard paradigm of a combinatorial mission problem right when you have a truck assessment we need to to find the optimal tour to visit a series of cities and so it has of course also direct applications in industry for example very vehicle routing right where a company has to deliver stuff and so it needs to find the optimal route and so uh yeah traveling sesame is certainly a great example and it also historically has driven most of the research in computer optimization and still it is also in the intersection with machine learning but also i'm thinking about the other problems like a cheap design or you know some problems combinatorial problems in wireless so these are also problems that are you know hard so in fact there are some mp complete you know problems there and uh you know these are also very interesting application areas for us so in chip design that would be things like routing traces on a circuit board or on a chip yeah so there are indeed a few stages of chip designs and all of them in fact involve different combinator optimization problems and uh yeah so as you say routing so finding the uh optimal connections between the different uh logical gates over for memories on a on a cheap canvas and also you you can think about indeed also placing these components on a chip in the optimal way to minimize area and stuff like that indeed this is all very interesting uh combinatorial problems where you know ai could disrupt and on the wireless side i'm imagining that's like frequency allocations and things like that or uh what are some of the applications there so i have in mind you know things like uh some coding problems where uh you know you basically have a sent some signal over and then you want to uh you know retrieve what was the uh uh b string that was sent but you know your scene has been corrupted by a noisy channel and you want to retrieve that that this string so that's a one example of things that one could do or some other uh error correcting problems and things like that yeah so there's a bit of a relationship between the combinatorics and kind of information theoretical types of problems and compression which is where you spend the bulk of your time yeah absolutely it is also fair to say indeed that uh all of these problems uh you know compression also other problems that we work on in qualcomm like quantization these are all of combinatorial nature so these are certainly possible use cases for this nascent field of ml for a combinatory optimization yeah nice nice yeah so your probabilistic numeric cnn's paper um what's the problem area that you're addressing there sure so there the uh the problem area is the application of deep learning to you know signals that are not necessarily sampled on a grid think about you know in many applications of deep learning you for example want to model uh um some time series and this time series you know have not been uh you know sampled uniformly so that's one more possible you know use case for uh the models we we want to develop like in probabilistic pneumatic cnns but more generally the the motivation there is really uh you know trying to to think about uh the signals that we we want to model you know in machine learning in their continuous formulation not in their discrete formulation which is you know the natural formulation you observe you know when when you measure something and so um basically uh just to unpack a little bit of the the title right so the title starts with the probabilistic numeric so the uh inspiration for this work is really this field of probabilistic numeric so you certainly are familiar with the convolutional neural network so i do not need to introduce that but so the pro let me just spend a couple of words on probabilistic numerics which might not be familiar to everyone so probabilistic numeric is a recent field in uh statistics which tries to um i mean not necessarily recent but you know recently i think there's been a quite uh big developments and um so it tries to quantify the uncertainty that a numerical program has due to the discreteness of the sampling procedure of its input to make this concrete let's think about the problem of computing a numerical integral you know the function that you want to integrate analytical you can evaluate it everywhere on a continuous range in its domain but necessarily you need to sample that function on a discrete sample of points because you know your memory and time to to compute that integral is finite and so probabilistic numeric tells you a way to uh derive uncertainty from the sampling procedure and the way it is done is by a bayesian inference so in bayesian inference one has a an agent a machine learning model that has a prior here the prior will be over the set of possible functions that you want to integrate and in technical terms it is a gaussian process it's a gaussian prior on this set of functions and then upon measuring that function you want to integrate on some points in your domain you update your prior to a posterior and this allows you to immediately translate to some uncertainty on the result of your of your integral so your now your numerical program will not just return a number with to return a probability distribution and this probability distribution will be picked around some value and there will be an uncertainty so that's your discretization error so we thought this is quite interesting also for machine learning and so we indeed start from the same philosophical standpoint you know the images that we want to model in machine learning the time series that we want to model are in fact continuous signal and so we necessarily need to discretize them because we want to put them in a computer and but you know this procedure will come with some uncertainty with some errors and uh so it is important to quantify those and so that's basically the motivation behind this work um if i could uh replay that to make sure i'm i'm understanding uh with classical numerical programs like i'm thinking back to fortran numerical computing and undergrad like you've got some function and you want to compute an integral for it and you just you do that there are established algorithms for doing that what probabilistic layers on top of that is allowing you to look at your the function that you're integrating not as a single function but as a distribution of functions uh and what you have then um you know then your kind of classic quantization error now becomes a probabilistic quantization error yeah that's a good summary thanks and maybe just to to clarify uh so here we are really looking at the quantization right in the domain of the of the function so the values that the function takes are still continuous and so that's the kind of quantization we are looking at and so indeed in our um you know probabilistic numeric uh cnn so we start from from this uh idea and then we uh develop on top a neural network i just wanna interrupt to say that uh when you say the the quantization is in the domain of the function meaning as opposed to the range which is your think about your vertical your amplitude here we're talking about you're taking different points in time that may or not be may or may not be well are not uniform and so that's where your quantization is coming in so you've got a time series but you're not getting data in every second or millisecond or whatever it comes uh irregularly and you're trying to figure out uh quantization error based on that irregularity that's precisely it yeah okay yeah all right cool so and uh indeed maybe just to give a little more about the the paper so we we develop the um you know idea uh of using probabilistic numeric for uh deep learning so um the first step in this procedure is that we start indeed from a regular sample time series for example or even from an image which uh you know has been sub-sampled in a regular way and what we do is that we interpolate that so we interpolate that in a probabilistic way so like exactly like you know probabilistic numerical programs do and that gives us you know a posterior distribution over our input and what we do then on this posterior distribution is that we apply a neural net so now this posterior distribution is a distribution of a continuous functions and so the technical contribution that we we make in this paper is to devise a neural network on continuous functions so typical typically your cnn will rely on vectors right so some array of numbers here our probabilistic omega cnn is defined directly in the continuum and that turns out to be quite powerful and also unlocks you know new models and a new mechanism for for learning whichever yeah what does that exactly mean um i i think yes i'm so used to thinking about the input to a cnn being a vector i'm not even sure how to how to unpack it being continuous right so indeed you're not gonna store that function in your computer because of uh by definition indeed you're gonna need to you know have uh you know uh an infinite number of points if you want to store all the values what you're gonna store is it's just some functional form some code that allows you to evaluate that function right and so that's about the input to your uh you know a neural network and so to be precise indeed about what happens we still have a neural network which works by interleaving linear non-linear layers but now and so the non-linear layer you can actually morally understand that it's going to be very similar to what you used to do at each point of your function you apply non-linearity but now the real you know new uh part of the work is about the linear layer so we devise actually a new convolutional layer which is defined in terms of a linear pd partial differential equation so this partial differential equation is is a linear operation on an input function which is the input uh sum out to the pde namely the value that you have the initial condition to your differential equation so um what happens is that uh you know uh if you want to impose actually translation equivariance that that you have you know in convolutional layer this restricts the forms of differential equations that you can put in your neural network and interestingly uh you know one of the simplest things you can do is to use the kind of generalized diffusion equation so diffusion you know is a process from physics which you can understand you know for example when you have a glass of water you put some dye into it and this dye diffuses over time and so similarly here you know we have our uh image which is encoded you know in some function and that function gets blurred similar to the diffusion process over time so that's really what what we mean you know by the layer on continuous function so it is defined formally and it turns out that for certain choices indeed of of layers we can do computations analytically so we can actually propagate this uh you know functional forms in our in our code analytically and and so that's a pretty cool um and so we can you know uh ultimately we we can devise a practical procedure to you know start from our input signal which which was you know this subsampled signal then interpolated then we applied this uh convolutional layer spds we interleave with some non-linearities and what we get out after some of these layers and perhaps the pooling and so on we get out you know a prediction like you know we want to classify this time series for example for this input image and so we want to get out the class label right as we do usually but on top of that we also get out the uncertainty and actually this uncertainty is also there at the every intermediate layer and it's really an uncertainty that is related to the you know fact that the input signal didn't have maybe information in certain regions of space or time so in this way we know you know we can characterize indeed what is the error that we make and more interesting we can also choose where to sample the signal in order to know to reduce uncertainty so these are all the interesting applications that we can think about with this model is that that latter point uh choosing where to sample is that um a like a byproduct of going through the process in the same direction or is it more like going through the process backwards i don't know if that question makes it's going to the process backwards your eyes somehow you uh you basically you know find a certain uh uncertainty and this uncertainty will be a function of where you value your input so you can compute some kind of derivative of that with respect to the inputs to minimize the uncertainty and that's that can be useful you know when it is for example costly to to get data points right you can optimize for the points that are most informative or you know when you have maybe some data meshes and things like that you know where discretization errors are important so there are a lot of interesting you know use cases um so in this paper actually we focus mostly on benchmarking this model on a couple of data sets so one is the super pixel uh classification of of images so it's super pic super pixel you know it's just an image which is sub sampled but again the the points are not on a grid before we get to the benchmarks i have another question about the architecture here so you one of the key innovations or contributions it sounds like is this pde layer and pdes arise in physics all the time like you can i'm imagining the inspiration of that was thinking about the problem like the closed forum problem and how you might solve it and then you know that involves pdes uh yeah yeah so big yeah go ahead oh no no i i was i was gonna you know but then you get to that so that you your pd you have this pde layer that you think needs to be involved in here but it you have the constraints of translation and variance from cnns and then suddenly you're like okay diffusion is the answer and like where did that come from was that did you did you uh recognize diffusion as like a translational independence by thinking about a glass of water or is that like a known physics thing or so yeah so the um the way we got there and actually i i i would like at this point to um amend one of the big omissions that i've done in the beginning which is not to acknowledge that the first author of this paper is mark finchy who was doing a an internship with us last summer so he's really the the main driving driving force in this project and so you know mark came up with this uh proposal and i guess it was a bit of a mixture of two things one thing was intuition and the other thing maybe coming from physical reasoning and the other thing was just mathematical formalism so we we wrote down you know the most general basically local you know linear layer in the form of the pde and then you know basically this turned out to be uh diffusion when you impose translation in variance and actually we also you know did something a bit more general so we also consider the you know symmetries like rotation and things like that yeah so we thought a little bit about spherical synapse at the beginning so there is a you know interest in the community in characterizing equivalence and the more general symmetries and so it turns out that you know beautifully also in this context we can get you know pds which are you know equivalent under more general transformations like rotations and things like that and that's actually quite interesting i believe because uh you know one of the problems with the you know getting to work also the equivalence and the rotation say is that you necessarily need to discretize things on a lattice and so at that point you know the rotation for a certain angle becomes uh you know pretty tricky to to get it to work well and necessarily you know you will have some error which is due to the basically mesh of your lattice and so on in this context we avoided this problem so our model is defining the continuum and you know it is basically equivalent under arbitrary rotations so that's a pretty cool i think uh feature and also equivalent on the arbitrary translation so that's i think a pretty cool feature too right yeah i'll just interject really quickly that the this whole idea of equivariance and spherical cnns and gauge equivalence is a big focus of the ai research team there at qualcomm and for folks that want to dig in more probably the best place to start is the first interview i did with max welling on gauge equivalent cnns where we talk about a lot of this uh what equivariance is and why it's important and we'll drop a link to that in the show notes yeah thanks i i also listen to that it was a great uh episode yeah awesome awesome so you were talking about benchmarking yeah indeed so um uh i i was talking about the fact that we benchmarked on on a couple of data sets so the first one was this uh super pixel images so you start from an image and you know suppose it is defined on a grid and then you subsample it so you you take away some of these points in such a way that then it becomes you know the degree structure is lost and you know at this point your regular cnn will not work well for this data type there are a few other competitors out there but it turns out that our model basically established a new state of the art for for this task so three times reduction in the the test uh error and so this was quite encouraging and we also applied the you know the uh model to a medical time series so in this case uh you know you you can think about you know a patient who goes to the hospital and then you know uh for example the the doctor measures you know blood pressure or things like that and this is done at irregular times right so this is also a good uh you know case of a irregular time series and then based on these measurements you want to predict you know if the patient will recover or things like that and in fact so this is our these are pretty important data sets to to look at and so we also apply our our model to these data sets and show competitive results there too yeah so um we uh we basically uh um you know think that uh this point of view is very powerful and um yeah in fact uh we have a few you know future directions uh in mind uh that came after this uh this work and what are those right so one of the interesting directions uh um for this work in my opinion is is to connect to to to quantum computation in fact um one of the uh promising platform for quantum computation is the so-called quantum optical computer so here uh optical means that you you use light as the you know um basically unit of information that you want to manipulate and it turns out that uh uh you know there are states of light that that you know people know how to to to build in a lab which are closely related to to gaussian processes so there is a beautiful connection here between states of light and gaussian processes and they immediately disperse a connection also between you know our model and um a possible general generalization to a quantum model so a quantum neural network and so i'm excited about this because uh you know it seems to me a very natural uh basically way to encode the data via this uh relationship between gaussian processes and certain states gaussian states of light so that's i think a very natural way to encode data in a quantum computer which operates on on uh with optical elements and therefore you know there is also a nice way to interpret uh basically the probabilistic numeric neural network from this point of view of quantum information and quantum optics the quantum mechanics so there is a big basically um direction here that i'm i'm excited about which spurred out of this uh of this paper actually and this new way to think about inputs to neural network and think about neural networks is the primary connection there thinking about continuous functions or is there also this aspect of uh missing or irregularly sampled data yeah i would say both so the the fact that the indeed you have continuous function uh relates to what you know physicists um call uh basically a kind of model that physically use which is called the quantum field theory so the quantified theory is also continuous field and so your continuous field classically right which is your function which is continuous becomes now a quantum function so it becomes a quantum field and this is quite exciting i think because this you know these quantum fields are relevant for describing you know elementary elementary particles so these kind of experiments that that you see in order in this big colliders like cern and so on so this you know continuous formulation allows you to make a very interesting connection between very different fields which are described by very similar formalism and so this is a very interesting thing and also the uh you know sampling uh the irregularly sampled nature of the inputs is also i think naturally captured in terms of of these gaussian states of light that i was talking about so yes i would say both in my opinion are very natural candidates you know that allow you to i think propose interesting models for for quantum neural networks and what are some of the you talked about kind of the the going back to the benchmarks the subsampled images and the um the healthcare time series um is there also a compression uh application for this paper as well no i would say that a compression was not our main focus but i can see maybe where you're going so you if if you can condense you know your input in some mean and covariance of a gaussian that's maybe also a way to think about you have compressed it to a few numbers so that's an interesting spin that i didn't think about so it was not really our our focus here but yeah it might be got it um cool so uh again this is a paper that you're presenting at iclr what else does qualcomm have going on at the conference sure so um another paper uh is from my uh colleagues this far rosendale and iris have been taco coin and um so this paper is about uh adaptive uh neural compression so here the uh so we thought a little bit before you know about what is the idea of neural compression so in your neural codecs and so typically there is a problem which is the problem that you know you want to have a small neural network because you want uh you know to uh to have a low computational burden to do this codec process but at the same time a low neuron low complexity neural network we will typically you know not generalize well right so and so the idea here of the authors in this paper that we present at iclr was to do adaptive or fine-tuned compression so the idea is that uh okay you have trained your model on some training data but now you deployed but it turns out that you know the test data can be different from the training data as i said that you can suffer from generalization problems but uh you can you know imagine now that uh in a scenario where for example your sender you know is some um you know at your sender side of the data you can you're willing to spend compute time and you know you're sending this data to this compressed data to some low power devices like mobile phones and in this scenario it makes sense you know that at your sender time you can fine-tune your your um basically codec on the test data and then on top of sending you know the transmitted image or video to your mobile phone you also send some bits that are related to the difference in the weights of your you know adapted neural network codec and so the authors indeed show that you know you can indeed reserve some some bandwidth for the for this uh delta in the weights and on top of the bandwidth for for the stream of the video that that you want to to send and it turns out that if you do that and if you jointly optimize the model to do the the best possible thing so to minimize the rate the number of bits if you transmit and opt maximize the reconstruction accuracy you actually can do better if you do this procedure you know sending over some of the bits free up daily weights then then if you just you know occupy the old bandwidth for your for your stream so that's a pretty all promising i think direction which can have a few interesting you know applications a direct application in platform also for qualcomm guys i think this one was counter-intuitive for me i i may be misremembering the information theory but i thought like nyquist or hammond or heming or something like that said that it doesn't matter how you chop up your your channel you know if you have a fixed bandwidth there's some certain maximum uh throughput that you'll be able to get but here you're kind of splitting your channel into uh kind of in-band and out-of-band or something like that and getting better results yeah so here the idea is really that you know you want to uh basically send a certain number of beats right so that's your somehow your uh the the rate that you are you're willing to send so that's somehow the uh amount of the information that you you you would like to send and at that point you would like to do you know the best possible job given the constraint so what is the you know uh choice of my codec from my uh you know encoder right to give me the best description of my input in such a way that when i decode it i get the highest reconstruction quality and so that's the the setup and so in this setup uh you know what we showed is that you can actually uh reserve some of these beasts that that you transmit for the you know weights and so that was a new idea that uh you know people have not thought about and uh yeah in this setting this is beneficial but you're right there are certainly some theoretical bounds to do to the you know uh rate distortion uh performance that you can achieve but you know we are certainly within these bounds uh and uh yeah um okay so yeah we're talking about the different the the thing that i was thinking about it applies to the theoretical bounds but it's not i wouldn't say that so we have some ability to operate uh within that envelope i would say anything closer to the bound with this procedure than without got it got it um cool any other uh papers uh qualcomm ai research paper is it iclr yes certainly there are a few other papers i can highlight a paper by my colleague pim de haan who is a phd student of max and he has a paper on mesh cnn so here the idea is that you know you have tasks or meshes like you know shape segmentation or 3d shape reconstruction and things like that and so um you would like to use a graph neural network for these tasks but graph neural networks suffer from the problem that they are they are oblivious to geometry so by definition of the graph structure they do not have the information about the geometry so in particular if you have you know a vertex with two edges connected to it you know it doesn't matter if you move these edges around basically because the graph neural networks do not conclusion neural networks do not see the angle between these edges and so the idea of this mesh cnn is to use you know gauge equivalence tools to build this geometry into graph neural networks it turns out that if you do that you get you know much better results for these tasks or meshes okay yeah um so we have also other works on um um you know uh temporal localization of actions uh so and um and also we have a um we are organizing together with the um uc irvine and uh disney research we are organizing a workshop at iclr work workshop on neural compression so that's certainly an exciting opportunity to you know get together with the experts in information theory communication and deep learning to indeed further advance this uh effort of getting better codex using neural networks nice nice um so you've talked uh we spoke earlier on kind of where you saw the probabilistic numeric research going kind of more broadly when you think about your research area and the the area of your team what what are you most excited about where do you see that going yeah so indeed we spoke earlier and i was hinting at a quantum neural network so this is certainly something that i believe would be a drive for innovation in ai in the future you know recent years last just last couple of years i've seen a tremendous you know excitement in the quantum computing community you know with the supremacy experiments so we are really at an era where we are starting you know to seriously think about uh you know these things um and so it's really timely i think to to to get serious about uh you know thinking about how can you use quantum information quantum computation to uh enhance uh machine learning so that's i think a very exciting place to be um it is true that uh you know it is still a recent field and um you know there's certainly a lot to do and uh it is still an open question to to figure out what's the the best way to use quantum computers for machine learning so that's why i think it's a very exciting area so i've been thinking did a little bit about this over the last year and um so one of the things also i i've been thinking about was the um the problem you know of of benchmarking this model so we talked a little bit about this direction where you can use the quantum optical computers and the relation to you know uh probabilistic numeric cnn and so on but you know in general the problem here is that we do not have these devices to run the quantum neural networks at scales that we would like right now right so what do we do certainly we can indeed based on intuition and small experiments figure out what are the most promising models and i think that uh i worked on was to also try to find an interpolation between you know your classical neural net and a quantum neural net so basically we came up with this quantum deformed neural networks um which is you know is some work uh i did also with maxwelli and um so the idea here is that um you know you you can think about your classical neural network as embedded in a quantum computer and in fact uh the architecture we are thinking about now is the qubit architecture which is an interesting relation you know to binary neural nets because you know a cubit is basically the you know quantum equivalent of a bit and so there is a relation with quantization of you know natural that you can also explore using quantum mechanics and so on but the uh the most interesting thing is that indeed we managed to map this uh binary neural nets or probabilistic binding uns in fact on a quantum computer using qubits and then we started to deform this uh this model so to introduce gates which you know use pure quantum effects like entanglement and superposition and so we we do that in a way that you know uh we slowly interpolate between a regime which we can stimulate classically which is basically the classical neural net regime and the regime we know which is basically uh intractable classical which will require a quantum computer and along the way there is a some some you know intermediate regime where you can still do some classical simulations using some tools from quantum thesis which are called tensor networks and basically this allows you to start from a good you know prior somehow for your model which is this classical network deform it by doing this indeed these modifications and then you can you know use classical simulations of the quantum model to learn that so that you can implement it as a differentiable program and so we show actually some modest gains with respect to you know the the classical neural net by introducing these quantum gates and we can actually provide you know the first example where you can simulate a quantum model at the scale of you know real world data so that was interesting for us but yeah more generally you know there are many in the different directions at the moment also related to optimization problems so we talked about combinator optimization right and machine learning and also you know quantum computing is also an exciting area for exploring new algorithm for combinator optimization so yeah to summarize indeed the quantum ai i think is a very interesting direction and also i also think that the machine learning for combinator optimization is a very interesting direction so this is also pretty recent and i believe that here we will see you know a large-scale adoption of this technique because it has been shown recently that you know you can actually enhance your classical solvers for you know the integer linear programs or stuff like that which you know people use routinely for solving their problems in the industry you can actually uh you know augment with neural networks these solvers in such a way that the decisions that these servers make are better informed and basically are faster and the idea here is that you know you can adapt basically your your uh solver to the data distribution that you really care to solve you know a a delivery company which will uh routinely solve you know the traveling assessment problem in the same city you know you do not want to deal with the most arbitrary hard instances of the traveling assessment problem and this is where machine learning i think will really put an edge so it will allow you you know to tailor your optimization algorithm to the instances that that you care about and ultimately and ultimately this you know leads to you know better performance for combinator optimization solvers and um yeah and also potentially you know it allows you to discover new strategies like using reinforcement learning as we have seen in alphago alpha fold so super excited i think awesome awesome well roberto thanks so much for sharing a bit about what you're working on uh and what some of the folks on your team are working on at iclr uh it's been really great chatting with you thank you sam pleasure for me too thank [Music] you

Original Description

Today we kick off our ICLR 2021 coverage joined by Roberto Bondesan, an AI Researcher at Qualcomm. In our conversation with Roberto, we explore his paper Probabilistic Numeric Convolutional Neural Networks, which represents features as Gaussian processes, providing a probabilistic description of discretization error. We discuss some of the other work the team at Qualcomm presented at the conference, including a paper called Adaptive Neural Compression, as well as work on Gauge Equivariant Mesh CNNs. Finally, we briefly discuss quantum deep learning, and what excites Roberto and his team about the future of their research in combinatorial optimization. The complete show notes for this episode can be found at https://twimlai.com/go/482 Subscribe: Apple Podcasts: https://tinyurl.com/twimlapplepodcast Spotify: https://tinyurl.com/twimlspotify Google Podcasts: https://podcasts.google.com/?feed=aHR0cHM6Ly90d2ltbGFpLmxpYnN5bi5jb20vcnNz RSS: https://twimlai.libsyn.com/rss Full episodes playlist: https://www.youtube.com/playlist?list=PLILZm3MRkvH83C46bZ4rPmB-jKWBltWkP Subscribe to our Youtube Channel: https://www.youtube.com/channel/UC7kjWIK1H8tfmFlzZO-wHMw?sub_confirmation=1 Podcast website: https://twimlai.com Sign up for our newsletter: https://twimlai.com/newsletter Check out our blog: https://twimlai.com/blog Follow us on Twitter: https://twitter.com/twimlai Follow us on Facebook: https://facebook.com/twimlai Follow us on Instagram: https://instagram.com/twimlai
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from The TWIML AI Podcast with Sam Charrington · The TWIML AI Podcast with Sam Charrington · 0 of 60

← Previous Next →
1 Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
Engineering Practical Machine Learning Systems with Xavier Amatriain - #3
The TWIML AI Podcast with Sam Charrington
2 How to Build Confidence as an ML Developer with Siraj Raval - #2
How to Build Confidence as an ML Developer with Siraj Raval - #2
The TWIML AI Podcast with Sam Charrington
3 Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
Open Source Data Science Masters, Hybrid AI, Algorithmic Ethics & More with Clare Corthell - #1
The TWIML AI Podcast with Sam Charrington
4 Interactive AI, Plus Improving ML Education with Charles Isbell - #4
Interactive AI, Plus Improving ML Education with Charles Isbell - #4
The TWIML AI Podcast with Sam Charrington
5 Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
Machine Learning for the Stars & Productizing AI with Joshua Bloom - #5
The TWIML AI Podcast with Sam Charrington
6 Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
Generating Labeled Training Data for Your ML/AI Models with Angie Hugeback - #6
The TWIML AI Podcast with Sam Charrington
7 Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
Explaining the Predictions of Machine Learning Models with Carlos Guestrin - #7
The TWIML AI Podcast with Sam Charrington
8 Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
Deep Learning: Modular in Theory, Inflexible in Practice with Diogo Almeida - #8
The TWIML AI Podcast with Sam Charrington
9 Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
Emotional AI: Teaching Computers Empathy with Pascale Fung - #9
The TWIML AI Podcast with Sam Charrington
10 Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
Statistics vs Semantics for Natural Language Processing with Francisco Webber - #10
The TWIML AI Podcast with Sam Charrington
11 Building AI Products with Hilary Mason - #11
Building AI Products with Hilary Mason - #11
The TWIML AI Podcast with Sam Charrington
12 Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
Reprogramming the Human Genome with AI, w/ Brendan Frey - #12
The TWIML AI Podcast with Sam Charrington
13 Understanding Deep Neural Networks with Dr. James McCaffery - #13
Understanding Deep Neural Networks with Dr. James McCaffery - #13
The TWIML AI Podcast with Sam Charrington
14 Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
Scaling Deep Learning: Systems Challenges & More with Shubho Sengupta - #14
The TWIML AI Podcast with Sam Charrington
15 Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
Domain Knowledge in Machine Learning Models for Sustainability with Stefano Ermon - #15
The TWIML AI Podcast with Sam Charrington
16 Machine Learning in Cybersecurity with Evan Wright - #16
Machine Learning in Cybersecurity with Evan Wright - #16
The TWIML AI Podcast with Sam Charrington
17 Interactive Machine Learning Systems with Alekh Agarwal - #17
Interactive Machine Learning Systems with Alekh Agarwal - #17
The TWIML AI Podcast with Sam Charrington
18 Location-Based Intelligence for Smarter Marketing with Klustera - #18
Location-Based Intelligence for Smarter Marketing with Klustera - #18
The TWIML AI Podcast with Sam Charrington
19 AI-Powered Customer Support with HelloVera - #18
AI-Powered Customer Support with HelloVera - #18
The TWIML AI Podcast with Sam Charrington
20 Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
Using AI to Simplify the Programming of Robots with Cambrian Intelligence - #18
The TWIML AI Podcast with Sam Charrington
21 Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
Increasing Efficiency of Healthcare Insurance Billing with NLP, w/ Behold.ai - #18
The TWIML AI Podcast with Sam Charrington
22 Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
Creating a Worldwide Financial Knowledge Graph with AlphaVertex - #18
The TWIML AI Podcast with Sam Charrington
23 From Particle Physics to Audio AI with Scott Stephenson - #19
From Particle Physics to Audio AI with Scott Stephenson - #19
The TWIML AI Podcast with Sam Charrington
24 Selling AI to the Enterprise with Kathryn Hume - #20
Selling AI to the Enterprise with Kathryn Hume - #20
The TWIML AI Podcast with Sam Charrington
25 Engineering the Future of AI with Ruchir Puri - #21
Engineering the Future of AI with Ruchir Puri - #21
The TWIML AI Podcast with Sam Charrington
26 Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
Deep Neural Nets for Visual Recognition with Matt Zeiler - #22
The TWIML AI Podcast with Sam Charrington
27 Introducing Psycholinguistics into AI with Dominique Simmons- #23
Introducing Psycholinguistics into AI with Dominique Simmons- #23
The TWIML AI Podcast with Sam Charrington
28 Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
Reinforcement Learning: The Next Frontier of Gaming with Danny Lange - #24
The TWIML AI Podcast with Sam Charrington
29 Offensive vs Defensive Data Science with Deep Varma - #25
Offensive vs Defensive Data Science with Deep Varma - #25
The TWIML AI Podcast with Sam Charrington
30 Global AI Trends with Ben Lorica - #26
Global AI Trends with Ben Lorica - #26
The TWIML AI Podcast with Sam Charrington
31 Intelligent Autonomous Robots with Ilia Baranov - #27
Intelligent Autonomous Robots with Ilia Baranov - #27
The TWIML AI Podcast with Sam Charrington
32 Reinforcement Learning Deep Dive with Pieter Abbeel  - #28
Reinforcement Learning Deep Dive with Pieter Abbeel - #28
The TWIML AI Podcast with Sam Charrington
33 Robotic Perception and Control with Chelsea Finn  - #29
Robotic Perception and Control with Chelsea Finn - #29
The TWIML AI Podcast with Sam Charrington
34 Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
Natural Language Understanding for Amazon Alexa with Zornitsa Kozareva - #30
The TWIML AI Podcast with Sam Charrington
35 The Power of Probabilistic Programming with Ben Vigoda - #33
The Power of Probabilistic Programming with Ben Vigoda - #33
The TWIML AI Podcast with Sam Charrington
36 Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
Intel Nervana Update + Productizing AI Research with Naveen Rao and Hanlin Tang - #31
The TWIML AI Podcast with Sam Charrington
37 Video Object Detection at Scale with Reza Zadeh - #34
Video Object Detection at Scale with Reza Zadeh - #34
The TWIML AI Podcast with Sam Charrington
38 Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
Enhancing Customer Experiences with Emotional AI, w/ Rana el Kaliouby - #35
The TWIML AI Podcast with Sam Charrington
39 Expressive AI-Generated Music With Google's Performance RNN with Doug Eck  - #32
Expressive AI-Generated Music With Google's Performance RNN with Doug Eck - #32
The TWIML AI Podcast with Sam Charrington
40 Smart Buildings & IoT with Yodit Stanton - #36
Smart Buildings & IoT with Yodit Stanton - #36
The TWIML AI Podcast with Sam Charrington
41 Deep Robotic Learning with Sergey Levine - #37
Deep Robotic Learning with Sergey Levine - #37
The TWIML AI Podcast with Sam Charrington
42 Deep Learning for Warehouse Operations with Calvin Seward - #38
Deep Learning for Warehouse Operations with Calvin Seward - #38
The TWIML AI Podcast with Sam Charrington
43 Cognitive Biases in Data Science with Drew Conway - #39
Cognitive Biases in Data Science with Drew Conway - #39
The TWIML AI Podcast with Sam Charrington
44 Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
Data Pipelines at Zymergen with Airflow, w/ Erin Shellman - #41
The TWIML AI Podcast with Sam Charrington
45 Web Scale Engineering for Machine Learning with Sharath Rao - #40
Web Scale Engineering for Machine Learning with Sharath Rao - #40
The TWIML AI Podcast with Sam Charrington
46 Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
Marrying Physics-Based and Data-Driven ML Models with Josh Bloom - #42
The TWIML AI Podcast with Sam Charrington
47 Machine Teaching for Better Machine Learning with Mark Hammond - #43
Machine Teaching for Better Machine Learning with Mark Hammond - #43
The TWIML AI Podcast with Sam Charrington
48 LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber  - #44
LSTMs, Plus a Deep Learning History Lesson with Jürgen Schmidhuber - #44
The TWIML AI Podcast with Sam Charrington
49 Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
Learning From Simulated & Unsupervised Images through Adversarial Training - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
50 Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
Jennifer Prendki Interview - Agile Machine Learning - TWiML Talk #46
The TWIML AI Podcast with Sam Charrington
51 Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
Evolutionary Algorithms in Machine Learning with Risto Miikkulainen - #47
The TWIML AI Podcast with Sam Charrington
52 Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online  Meetup
Learning Long-Term Dependencies with Gradient Descent is Difficult - TWiML Online Meetup
The TWIML AI Podcast with Sam Charrington
53 Word2Vec & Friends with Bruno Gonçalves -#48
Word2Vec & Friends with Bruno Gonçalves -#48
The TWIML AI Podcast with Sam Charrington
54 Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan  - #49
Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan - #49
The TWIML AI Podcast with Sam Charrington
55 Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
The TWIML AI Podcast with Sam Charrington
56 Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
Intel Nervana DevCloud with Naveen Rao & Scott Apeland - #51
The TWIML AI Podcast with Sam Charrington
57 AI-Powered Conversational Interfaces with Paul Tepper - #52
AI-Powered Conversational Interfaces with Paul Tepper - #52
The TWIML AI Podcast with Sam Charrington
58 Topological Data Analysis with Gunnar Carlsson - #53
Topological Data Analysis with Gunnar Carlsson - #53
The TWIML AI Podcast with Sam Charrington
59 ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
The TWIML AI Podcast with Sam Charrington
60 Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
Ray:A Distributed Computing Platform for Reinforcement Learning with Ion Stoica -#55
The TWIML AI Podcast with Sam Charrington

This podcast episode explores the concept of Probabilistic Numeric Convolutional Neural Networks, discussing how Gaussian processes can provide a probabilistic description of discretization error, and touching on other research topics like Adaptive Neural Compression and quantum deep learning.

Key Takeaways
  1. Understand the basics of convolutional neural networks
  2. Learn about Gaussian processes and probabilistic modeling
  3. Implement Probabilistic Numeric CNNs
  4. Analyze discretization error in CNNs
  5. Explore other research topics like Adaptive Neural Compression and Gauge Equivariant Mesh CNNs
💡 Probabilistic Numeric CNNs can provide a more accurate and robust description of discretization error, leading to improved performance in certain applications.

Related Reads

Up next
RNNs Explained in 60 Seconds #ai #coding #machinelearning
Ascent
Watch →