PyTorch 2.0 Q&A: Dynamic Shapes and Calculating Maximum Batch Size

PyTorch · Intermediate ·🧬 Deep Learning ·3y ago

Key Takeaways

Explains dynamic shapes and calculating maximum batch size in PyTorch 2.0.

Full Transcript

hey everybody welcome back to the products 2.0 as the engineer series welcome back if you've been following along welcome if this is the first episode uh this is the series where you can ask questions to our excellent Engineers here on some of the new capabilities in pythons 2.0 and this is an ongoing series we have had several of these uh you can find uh all of these in the pytots 2.0 pythons.org events uh events page there you go I can't type today um events page you'll find all the upcoming ones and also the past ones and if you're watching this on YouTube you know where you can find the previous recordings they're all on YouTube feel free to go check those out and uh uh let's start with a quick round of introductions and uh we will get dive deep into the topic so I am Shashank I'm a developer Advocate at meta and uh who wants to go first sure both girl I'll go ahead so hi I'm Edward uh I've been working on the pie church project for uh five years now um You may also know me because I also do the pie Church Dev podcast although I haven't done an episode recently but uh hopefully hopefully we'll be starting it up again soon and um I also have a YouTube channel where we um have posted other stuff so um you know happy to answer questions with uh anyone who's got them that's awesome good uh you're muted I think the joys of live stream um if you're if you're listening to us uh tap something in the comments uh I love you yeah yeah so yeah let us know uh if you have questions that you want to start asking right away uh we'll we'll start uh we'll start the show with questions uh we love questions and that's the whole part of the series a quick reminder to keep asking questions this is not uh brought it's not a webinar this is supposed to be a live q a so we will interrupt our presentation to answer your questions so please ask questions um can you hear me now yes we can yes I've been an engineer on uh pie George for a while and uh I've been working on pt2 um for you know as uh in the last year or so that's awesome uh welcome Alliance welcome Edward we even have uh um uh fight out celebrity if you will right with Edward here so um today we are talking about uh Dynamic shapes right and when I think of dynamic shapes I think of you know variable and sequences or variable sized images or even bad sizes I suppose is that the right um topic we're discussing and we already support this in pytorch uh what's new what do we what do we have in pythons 2 that's going to be different um so whoever wants to take that first and let's kick off the discussion there I guess that's right so shashanka mentioned by the way thanks everyone for commenting Aki Sandeep Charles Sai and Vincent great to have you on the stream hope you can ask some questions uh so if there's a general question about pd2 we'll get to that later um so we're here to talk about Dynamic shapes as Shashank said Dynamic shapes are all about like you know Things That Vary over time right like if you have a bunch of things you want to do training on but like sometimes you've got 10 things in your batch sometimes five or you know you're doing training over sentences and you know sometimes sentences are long sometimes they are short that's you know exactly what's going on uh when we're talking about Dynamic shapes so if you're all about e or Pi torch there's really nothing more to say like pytorch works if you change the sizes it's because we you know put in a bunch of kernels that uh work uh you know no matter what kind of size input you give to them and that's that's it that's the end of the story but with pie Church 2 we have a compiler and when you have a compiler it's not so easy to figure that out because you know we want to do specialization we want to like give you the fastest code that you know can possibly run for any given operation you want to do but uh you know sometimes you're like well but I don't want to only be able to handle sentences that have hundred characters in them so you have to do something different in that case um Elias anything you want to add here um nope that sounds good thanks so you mentioned Edward you mentioned performance and partosto and I I see a question that's about what's interesting and Point Let's do so the way I see the broad overarching theme there is performance with by touch 2 right with torch uh inductor yeah that's right and there's a in-depth get started guide which I'm going to share right now in the live stream for those interested in checking out what's new in Python and today's topic is related to that it's one such topic and we've been doing live streams on all the individual topics so uh the two links I shared with you the blog post and the events uh is is your go-to resource right now one thing that I want to add is that pie turns 2 isn't just like we're going to do 2.0 and that's it right Putters 2 is the beginning of pie torch on compilers right so we've got Dynamo which is the new thing that lets you actually you know compile your code this time we've got it right because sometimes you're like well didn't you do that with tour script and the answer is well yeah but dorscript didn't work so well but uh you know this time I I really think we've got something good and so there's gonna be lots and lots of things that are going to get better over the future now that you have this new fundamental capability and dynamic shapes is one of those things because uh well as you will well I'm having to answer questions about it but you know some stuff works some stuff doesn't work and we're working hard to make it better and better as things go along okay so when I think of uh Edward Dynamic shapes I almost think like compiler is sort of incompatible with variable shape things am I thinking about it is it it's kind of true or it's intuitive and I just don't get it in principle uh by the way Sandy yes the these will all be recorded and you can view them on the series uh Shashank with you right now uh thanks for asking uh all the previous recordings are on YouTube if you're watching this on YouTube if you're watching this on LinkedIn head over to YouTube after the stream um and you'll find the previous recordings there and I'll share the link so to answer the question about like is are all other compilers static the answer is kind of yes um I mean in principle it should be possible to write a compiler that can handle Dynamic shapes and like if you think about like a c compiler or C plus plus compiler right like obviously you know they they can handle all this kind of thing not a problem but in deep learning you know there's a very very heavy bias towards uh you know having compilers that only work with seti shapes so we also think that you know hydrog2s support for diamond shapes is kind of new and interesting it's not something that I mean there's plenty of work going on in other projects to like bring Dynamic shapes but you know we we like have this as one of the really important things we wanted to uh have in the 2.0 release because you know pytorch's history as an eager mode framework means lots and lots of people you know love to actually write code that's very Dynamic and that's you know that's the thing thank you we have a lull on the questions so Elias has prepared some stuff for us to sort of uh get the juices flowing but as I as we said this is not a webinar we're just here to answer questions but Elias you wanna you wanna show us what you got um sure um so uh shank do you want to share my um screen I suppose yeah um so I've shared this uh collab um on uh fake tensors and and torch dispatch which are some of the uh mechanisms behind our our Dynamic shape support um and they're also kind of useful um you know in their own right so we're gonna walk through it and uh you know ask questions as they as they come um so first thing just make sure you have your build right um this is available on the nightly but um you know you have to make sure uh you're recent enough so um anyways torch dispatch we're going to talk about fake tensor is um quickly calculating shape mismatches in your code without any side effects then uh finally you know training with the maximum batch size quickly this is a question we get every live stream so I just want to mention so everything with uh in this live stream that we're showing are are in the night list they're preview they're in the night list so if you want to get them I think we'll share this notebook with you it's on the top so to go get the nightly Docker container or pip install from the night at least for you to try these out they're not available on the One X right so I mean that's obvious so I'll share the link to uh this document I know there's a question here on uh can you find access to the code and I'm gonna share this notebook right now sorry go ahead um thanks um so what is torch dispatch uh so torch dispatch was a way for you to overload Pi torch code with your own custom implementation um kind of in the same way that you can run model code with float32 or float16 you can now extend it with custom data types or or semantics interpret things lazily or really however you see fit there's a more detailed explainer that I've Linked In the collab horses explainer here so if you you know want all the Nitty Gritty I'd recommend reading that um at a high level you know there are tensor subclasses in their uh torch dispatch modes um so we're gonna walk through some examples and and look at them so this is just a very simple function we're doing a map mode and then we're um adding it with a random value a random tensor um so the first example of um overwriting this with custom functionality is through a subclass um so we have this logging tensor which all it's going to do is print out the kernel that we're running as we're running it um and there's a little bit of boilerplate in the new and um this uh torch dispatch method is actually where we override the functionality and the function here is the kernel that we're going to run and then this right here return function are quarks is uh where we're actually running um the function uh no dispatches is um disabling um the torch dispatch so we don't like infinitely recur um so here we're creating a logging tensor from a tensor and then um we're running it and kind of as expected we as we're running it we print out the function so you see that we're running a mammal and a ad um and then there's another way of doing the same functionality which is as a tensor mode um and this is kind of similar um except there's one global mode as opposed to individual tensors um being a subclass themselves so here a little less boilerplate we're printing the function you know with the kernel here and the arcs and arguments to the kernel um so the difference here is uh Rand is captured in this Trace because um modes also interpose on Constructors where you know otherwise torch.brand would be constructed in um this torch ibran wouldn't be a logging tensor unless there's a global mode installed so and it also captures the backwards so as you can see we also have these kernels that are run as part of the backwards so this um and then the other map moles um so that's a little intro into uh torch dispatch again read the the longer post um if you're interested in more details um but so uh that gets us to fake tensor so fake tensors um are kind of tensors that don't actually have any data attached to them but they have all of the same properties that they would if they did have data so you know the shape the strides the uh device those are all captured in the fake tensor so they kind of they allow you to run your code and reason about the shapes your program would have had or or other properties without actually running the kernels or you know requiring the particular devices or accelerators um so if you're familiar with Meta Meta tensors which uh had worked on um sort of a similar thing um only fake tensor is also um record the device um the uranium model so you know which parts of your model were in CPU which one parts were in uh Cuda or other devices and we use them a lot throughout 2.0 to trace through pytorch um and uh to kind of Reason about your program and do optimizations um so one use case is you know you we kind of have this long-running issue to statically text type shapes and you know that's very difficult because Python's very hard to statically type anyways and then tensor programs are also um can be very dynamic as as Ed was referring to earlier and uh hard to type in their own right so you know maybe who knows when we'll eventually get to the statically checked shapes um we're certainly not there yet but um one thing you can do is is run with fake tensors to like you know check if you have any shape mismatches without doing you know slow kernels or like actually touching your your data um so here we run we install the fake tensor mode which is a tortistat dispatch mode um and then we run this uh model and take the back where word um and uh no it's not actually faster than running kernels yet but um you know that's one thing we're tracking and hopefully uh soon we'll optimize it to run a little more quickly um is a new capability right it's that's new with 2.0 features too right or is it something that we've had in the past um so fake tensors were a part of um a like torch this x like kind of experimental sub Library okay and now they've like been re-implemented in core and um are kind of more tightly integrated with the pt2 stack and are in python and a little bit more um hackable um so some of you may be familiar with its previous iteration but it's like more widely available um now awesome so here you know we we're running our model um with fake tensors um and we get this error message that says we need to have you know channel three channels instead of one um we rerun it with the right input size and we don't get an error so um you know that's one you could use cases just um checking if you have any like logic errors in your program um so another thing we can do is um Computing um the maximum batch size um so uh to do that we're going to have to keep track of all of the live fake tensors in our program and like this the storage sizes um that they you know would take if they were real um and to do this we can use another torch dispatch mode they're composable to compute maximum live memory um and just one thing one kind of disclaimer is that sometimes in practice uh the maximum batch size may actually be less than what we're calculating here um because of you know memory fragmentation in the allocator or other sort of uses of memory but this is like the ideal um you know memory use that we're going to calculate a quick question if you don't mind uh alive uh more I guess fundamental knife question is why why would we want to calculate the maximum batch price it's mostly performance related right how much can you shove into a device um yeah is that the main motivation for why we would want to go through this exercise um yeah it's kind of it's performance related yeah you like want to shove as much training um data in one run as you can um without you know running out of memory to give a little more color on this like the reason why you want more data is twofold so one is that there's fixed overhead to like launching kernels on Kudos so if you like put a teeny tiny amount of data then you're gonna mostly be overhead bound and well Piers 2 is here to like help you be less overhead Bound by fusing kernels together but like in general like the more data you shove through the less and less framework overhead matters so bad large batch size helps there but there's another reason as well which is that like it's actually different to train a model on a large batch size than a small batch size because some operations will actually normalize over all the elements in the batch so the bigger your batch is the you know more like examples you can sort of average over when you're doing one of these reductions and in many cases that also helps convergence time uh we got some questions in the chat okay the the letter to what you said and then we'll start picking some of those questions uh you said it's different did you mean that mathematically different mathematically different so like if you train them all it's batch size 8 or 16 like you'll get different results in the end even okay it may or may not be desirable but it's definitely desirable from a performance standpoint yes okay okay let's do questions uh Elias you want to answer the fake tensor one yeah so you can do these things without fake tensors what is the benefit of using fake tensors um so I guess a few things um one is like it can be kind of annoying just in Cuda when you run it and um you get the um you have to reset everything and it's like just a little bit of an annoying process not that big of a deal um but another thing is um by by using fake tensors you can like more easily say oh like how does my memory use scale as my batch size increases even beyond the like capabilities of your device or um put some of your fake to put some of your network on CPU um like you do checkpointing or like checkpoint activations to to CPU and um kind of just reason without the constraints that your actual device has um so kind of a um yeah the dreaded um um kind of um one one use case that uh torch this x um uses fake tensors for is uh deferred initialization which is where you have some huge model and you kind of want to Shard it over many um different uh servers or devices and you're kind of initializing with fake tensor tensors gives you a way to reason about uh where that sharding should be without you know blowing up your one server that can't hold the big model um so it's a way of reasoning about your program without kind of the constraints of the device you definitely could do a lot of the stuff with actual real tensors as well in fact before we had fake tensors we did all of pie Church 2's compilation with real tensors and it mostly worked except you know when the compiler would oom while it was compiling because it was trying to like shove the entire model and actually executed the fake tensor helped a lot with that I really want to answer JD's question because it's a really interesting question so JD's question is for dynamic shape in torch RNN pack padded sequence and Pad pack sequence are very useful but it has issues with Onyx export is there a good way to handle such cases so unfortunately I don't have a direct answer for how to help you with your Onyx export because um uh I haven't like done very much stuff with Onyx for years but um pack padded sequence and Pad pack sequence are very interesting functions because um they are what we call data dependent operations so um to explain what a data dependent operation is um let's think about a simpler operation so torch.nonzero what does this function do so torch.zero says give me a tensor and then give me a new tensor which only has the non-zero elements so you know however many non-zero elements in it that's what you're going to get so how big is the output tensor you get in this situation well you might think well you don't know right because if there's 20 you know non-zero elements you'll go to size 20 tensor if there are five you'll get a size five one so that's what I mean by the output shape is uh data dependent and so this traditionally causes a lot of problems for you know all those systems that are like oh everything is static and they're you know no problems well uh oh you've got this tensor in the middle of your network that might vary size over time so what are you going to do well one of the things we worked really hard on um when adding Dynamic shape support to Pi Church 2.0 was like actually being able to handle this case so in these particular cases um what we do today and what we plan to do in the future are different so what we do today is um if you want to uh like torch.compile one of these things so not an onyx textbook but torch.compile we are going to do what's called a graph break and so this is very different right so if you're exporting a model um you know you need it all in one piece but part of what makes pie Church 2 work on so many models whereas with tort script you had to just you know you had a like torque scriptify your models it didn't really work out of the box is whenever we see something we don't like we're just like okay we're going to start compiling we're going to go back to error mode and then we'll pick up compiling again after we've gotten past the operator that's being bad and terrible so we get to non-zero we're like uh oh this is a data dependent output all we're in a graph break and then we have a new graph which takes in whatever the output of non-zero was and now here's where Dynamic shape support comes in right if your compiler is able to compile the uh the suffix of your computation the thing after non-zero with Dynamic shapes then you are fine right because like a non-zero gives you 20 gives you 40 whatever you've got a single kernel that works in all those cases and you don't recompile so this is what we expect to work today what we want to do in the future is we want to be able to capture this entirely in one graph without having a graph break and now that helps you out if you're trying to export like say to Onyx because you can actually get a graph with a non-zero in it and you know everything is great and to handle this um you know uh one of the things that I've been working on is this thing called unbacked symbolic integers that's a mouthful so let's unpack it right so why are we talking about integers is because we're talking about the sizes of tensors and those are integers natural numbers really but you know we just call it editors symbolic because we don't know what the integers are right we want to represent them symbolically without reference to a real value like when you call non-zero you don't know if it's a five or a ten so we're going to call it s0 we're going to say well I don't know what it is but it's s0 and unbacked means that I don't know what the actual value is for a lot of things we do with Shake computation in pytorch we know what the symbolic variables are and so if you say ask hey is this shape equal to 20 we can look at what the actual underlying thing is and that'll tell us whether or not it's true or false but for a unback similar integer we don't know it looks like we lost a Lotus maybe we'll actually back in a moment we don't know what it is and we have to actually go ahead and um we have to uh we have to reason entirely symbolically and so if you want to do this conditional that doesn't work we have to give an error in this case the rewinding back to the question right so how do we actually export pack pattern sequence well these operators are also data dependent because what you're doing is you've got this sequence which is padded and you want to pack it into a sequence that's not padded how long the unpadded sequences depends on you know how much padding you actually had in your tensor which is a data dependent thing and so if unback Simmons work and dynamic shapes work then you can actually export these programs without actually you know needing to uh you know like it just doesn't work today with Onyx export and with Dynamic shapes it will work well that's the hope we're not quite there yet yeah I think Edward you you answered a lot of different questions in that in that explains it's a great question I really want to answer that question yeah there's so much some of the things I took away was your your explanation around uh the graph break thing that makes total sense so if you had these um uh these Ops that you you couldn't you have to have two separate sub graphs right because you had some Dynamic shape thing here then you couldn't do operator Fusion across it you could do operational fusions before you can do it after what you could do you couldn't do it because that is there right and if you have many of those then there are many such breaks and as we as you discussed in the beginning that will mean more overhead right I get that summary right more overhead for device Communications and and all that um awesome thank you for that that was that was a good explanation uh I while we wait for questions there are two that are sort of a digression to our topic I don't know if you want to take them feel free to take them this is about support yeah on Apple silicon two the two people who ask the question so I'll just summarize that is is it supported um and and related is mixed Precision supported yeah so we at at the current time and for the upcoming 2.0 release we're not going to have a MPS backhand for pi torch 2.0 however there's no reason we couldn't have unraveled it a little bit for the rest of mpss oh yes MPS is Shader okay that's your equivalent to I guess not Coda but one level lower for a GPU user so that's sort of the um well Coda is a level higher than the GPS MPS is kind of um it's higher level than Cuda in life oh it is okay okay yeah because but it doesn't really matter it's just this is the thing for Apple gpus so if you want Nvidia you do Cuda and if you want Apple it's metal or MPS or so that's a low level API to program a g uh an apple silicon that's yeah that's right so it totally makes sense to do it it's just like sort of not currently announced so far oh guys and uh next position should work without with MPS not on 2.0 it does yeah integer mode and also if you want to just see if you do CPU training on Apple silicon with mixed Precision that all works and like I mean you know the the the Apple silicon process CPU processor is also pretty nifty and we have the fine folks at well weirdly enough Intel um working on the CPU back end in Fighters too and they're they're also hard at work um you know giving you good CPU performance as well so check that out as well that's awesome uh end of the day I think the developer experience should be a tossed out compiler and everything is magically faster on every Hardware that you're running this on I guess that's the Holy Grail user that's the Holy Grail yeah there was also a question about um on TPU and there actually I have more to report so the way you can run pytorch on tpus uh today on tpus being Google's um you know deep learning Hardware um is via our pytorch xla integration which uh you know basically it's kind of it's kind of like uh pre PT 2.0 where what they do is you run your model and you capture all the operations that happened on it using what's called lazy tensor and then they ship that off to exlay to compile and so they are actively looking at integrating directly with Dynamo for xla so yes Fighters 2 and xla and tpus this this is this is this is a go this will happen in the near future yeah I think we had a talk on um uh feel free to check it out for folks who are joining in today uh we had a talk on back-end integration uh where we where we discussed um how new how Hardware went as like you know tpus or um gpus and others could actually um integrate with pytos 2.0 graph compilation system right the torch Dynamo and torch um inductor and all that so there's a set of low level apis they can use to integrate which is different from the previous path which was uh xla which brings me to a question that I have and I don't I don't know if it makes sense in this contest but would the the symbolic support that you're doing with Dynamic shapes how will it percolate down to the I guess torch Prim and further down the stack for Hardware would they have to each address it differently I I don't know if that makes sense how would that flow through the existing compiler system in 2.0 did my question make sense yeah yeah so [Music] I guess like with prims or with Dynamic shapes like it's it's gonna be infrastructure that's like available for them so in the prints you know kind of limiting The Operators they'll have to implement for pi torch which makes it easier to add a back end um and for dynamic shapes it should help them you know like consume like uh what is how is this you know size calculated from the inputs Etc so um hopefully you know the kind of information that we capture at the higher level will be used by back ends and in kind of the same way that we've used it for our own back end um as things slow down the stack awesome uh Andy Rock has a question yeah yeah is symbolic rank a planned feature for dynamic shapes i.eufunk style arbitrary leading bash Dimensions so uh just to um help the others out on the Stream So what Andy means when he talks about you Funk is this concept in numpy called uh youfunk shirt for Universal functions and it's a very fancy way of saying basically things that are like admull or uh div or some of the things that you know uh their point-wise operations and you can add as many dimensions as you want and they sort of generalize in the normal way um so in this particular case um a symbolic rank might refer to like hey I've got a tensor and maybe I want to have one batch Dimension or two batch or three you know as many as you want uh you know why not uh whatever and so unfortunately the answer to the question is no no symbolic ranks for you if you really need this to work what you should do is you should reshape your tensor to squash all the batch Dimensions into a single Dimension and then pass that to you know your torch compile thing the reason why we're not doing symbolic ranks is when it's complicated because you no longer get to do reasoning only on symbolic integers remember I was talking about symbolic integers right like this is sort of the lifeblood of dynamic shapes we just treat we have a bunch of integers and we do computation on them and we treat it symbolically and that's how we actually compile it if you have symbolic rank it's not just symbolic integers you have a like symbolic list which might vary in size and Elias has some experience with this because uh he spent some time working with Z3 formulas for uh you know doing shape computation yeah it uh it it makes everything a lot more complicated and like for our particular implementation and the new um uh new Dynamic shapes like it means you can't like just kind of Trace through operations and record them like you now have higher order operators that you'll need to contain like looping through Dynamic lists things like that um but as as uh Edward mentioned um like using reshape or using a torch dispatch subclass potentially to uh kind of contain your your higher level multi um uh unbounded batch Dimensions um tensor and and actually implement it with with fixed Dimension um A10 operators might be a Way Forward which um you know relates to the the collab I'm sharing we don't think there are too many use cases where you know you really legitimately need uh symbolic ranks um the one that I can think of that isn't uh well okay so there's two I can think of so one is um uh pyro the probabilistic and some other probabilistic programming Frameworks they often like allocate a dimension per variable that is being reasoned about so you'll get these like wacky like 128 dim tensors and like that's going to work very poorly with torchic.confile because we don't have symbolic dims sorry so that won't work well and the other thing I can think of is if you're just trying to like compile a kernel that like yeah is general purpose and doesn't recompile um you know like a you Funk right you want to write your own you Funk and then compile it with torch compile and then call it in a normal Partridge ear program and the answer to that is no we just can't do that you'll just end up with a specialization for dim there's a non-dynamic shapes related question but since we don't have any other questions um we should do it uh yeah that is a new question and uh I'll just read it out so uh we have people joining in from different platforms so I'll just read out the questions so everyone know so this question is about as pytos 2.0 became more complex it became harder for newcomers who want to dive into low level components to understand it what would you suggest as a resource for better understanding uh do any of you want to take that Ian we can each give our takes but um I think uh it's true it's it's more complex there's a lot of different components um I would uh like hopefully these developer series other kind of posts that we've been making will be a good enough introduction to any one of those components whether it's like uh you know dynamic shapes as Edward is mentioning or like our autograd handling or or our lower level or python tracing or our like lower level Cogen um so hopefully they're good resources and tutorials there um I'll start there um you know go to the code base like look at recent PR's look for um small issues um at the end of the day you know the best way to kind of understand something I think is is to contribute so looking for open issues or recently kind of clothes smaller ones and reading code um I guess would be my my thoughts but I don't know Ed what do you think well in talking about contributions this is one of the things where I think piders 2 is actually substantially better or than some of the other sort of compiler uh you know attempts that we made I'm I'm thinking specifically of tourist script and the reason it is better is because all nearly all of Pi Church 2's stack is written in Python Dynamo the thing that does the byte code all in Python inductor our compiler which compiles to Triton for getting kernels all in Python so so you actually like you don't have to like you know be a you know like build environment Rockstar and have a local build of Pi Torch from source to like actually get started hacking on any of this stuff you can like literally just go and edit the python files or like look at the back traces that python normally gives you yes they're very long back traces and you kind of need to know what you're looking for so yes it's more complicated because like you want to work on Dynamo you need to know what's going on with byte code you know what what symbolic evaluation you want to work on inductor you need to know a little bit about Charter and you need to know about uh you know how exactly are to find my own code generation works so yes there are more Concepts to understand but like if you just like like tinkering around and like changing things adding prints seeing what's going on it's way way more accessible and we're going to keep this stack in Python like basically as much as we can um we're actually like dealing with some Face Time Performance issues where it's like slower to run fake tensor than it is to run CPU because CPU is all in C plus plus and optimized because eager mode required it fake tensor we did it all in Python and so it's not very optimized and it's slow but like there's ways to optimize it which don't involve writing C plus plus and like that's sort of how we're pushing on this stuff the uh the only thing I want to add to that is you didn't mention what um role you're going to be playing so I kind of see three roles right you have the user which who ideally don't doesn't have to care about all these things right with tors dot compile things to work right from a user experience point of view it's straightforward or you're sort of a back-end integrator you're you're a hardware vendor or you're building something at the low end then you have the backend integration apis and if you're somewhere in between then I don't know what your role is other than you're a contributor uh or a maintainer you want to contribute or maintain and in which case the resources are like the dev dot Dev discus is an amazing place there's a it's a gold mine of information and you can ask questions and there's also the contributor slack uh where you can ask questions so if you're not at the vendor integration or the user somewhere in between you're trying to learn I think those are the resources in addition to Once uh Edward mentioned just contributing is sort of one of the best ways to learn the intermediate part of the uh stack um yeah there's there are two more questions um I I I I'll ask it now what about the thought script and the future of task script I guess to put it put it uh the extent we feel comfortable discussing right without yeah I think it's um safe to say that you know we're we're not super actively making contributions there I don't know about exact deprecation um lifetime so uh yeah I mean certainly for um like training or or you know whatever we recommend using torch stock compile um for some of the other use cases um like such as server Edge yeah there's no not a not a firm deadline yet on replacements okay we do have fans of that's good uh in the in the comments so uh there's another other question about when when will there be a launch of 2.0 Alpha I I don't know if you use those designations but it's in the night list today uh if you want to try it out and the ga will be in the future uh I don't know if you're giving we have any specific dates but do we use the alpha designation or I I think we're like technically an alpha right now and then when we that's what I released 2.0 like torch dot compile is going to be a beta feature I don't know but like basically like you know Intrepid users are testing it out today and sending them tons of bugs thank you for the bugs um we're working hard at fixing them there are a lot of folks and um when the release comes around we're hoping we'll have something useful it'll be buggy um like no no doubt but it will work on a lot of models and incremental updates yeah and we will be working hard on like fixing problems as it goes yeah uh and yeah I can call out please try it it's in the night at least like um I I just start by running the examples that's on the blog post and then take it from there and of course please let us know um that's how we improve of course it should be buggy in in loud ways hopefully as well so if it's buggy you'll know is that uh here's another question how to remove keys from dot State default pop method doesn't work I don't know if it's related to our topic of discussion unless uh you still want to take take it no idea but I like paz's suggestion use Dell uh that that makes sense to me too there we have a community we're building a community oh actually oh so maybe the problem is that uh uh no so Dell won't work either because State dick returns a shallow copy of the dictionary in question [Music] um probably what I would do is I would just look at the source code find the underlying attribute that actually has a thing and delete that I don't think we have a public API for this sorry okay we've got an actual Dynamic shares question and I won't answer it is there a document for dynamic shapes yet thank you Amir so yes there is so um so I've been posting weekly updates on the status of dynamic shapes on dev discuss so if you're like really interested in the most up-to-date info about what's going on you can check that out but I've also got I just quickly interrupt you just to say that this question came from LinkedIn and uh we are posting links but it's not going to LinkedIn it's only going to YouTube uh what we'll do is at the end of the stream I'll post all those links on LinkedIn but if you want it right away because you're you're uh impatient then they are available on the YouTube stream so sorry for the glitch but uh just want to mention I am posting links but it's not showing up on LinkedIn showing up on YouTube sorry continue Android the other thing that I'll add is that we've also got a dynamic shapes manual which I've shared on the on the dev discuss posts and this manual is essentially um uh it's oriented towards people who are like working on making their code work with Dynamic shapes because Dynamic shape support is sort of this ongoing process where some operations in pyter support Dynamic shapes and some don't and the manual tells you how to like go ahead and like extend support for dynamic shapes to other operators and this is especially important for us internally like internally we use pytorch and we have a lot of custom operators that like sort of aren't in open source because they're just like random stuff that like you know people play around with and like have unstable apis and stuff like that and they're like hey we want to use torreshold compile all this stuff and we're like okay but you got to write a what's called meta function for the operator saying how to actually do the compute and how do you do that well you can check out the manual to find out that um if you're more of an end user and you're just curious about what's going on with Dynamic shapes like how do I tell if it did anything or not there is a case study uh for open nmt which is one of the more recent posts in the thread that's also worth checking out we can look at that in more detail if uh no one asks more questions awesome um yeah I guess the two main things that the manual and the blog post the blog post um it's on dev discus just to clarify for everyone which is the discussion forum and that's a gold mine every time I go there I find I learn something so it's a gold mine of information that uh don't go to the pytouch blog uh you won't find it there but it's in that discuss um all right so did we cover contribute how um people can contribute um or do you wanna I know you touched upon it do we want to spend a couple more minutes just showing the ways people can contribute to this if they want to uh no I'd rather talk about like what this inductor code looks like with Dynamics okay um we don't have any more questions are there are there uh any I guess uh key takeaways that we want to share or what we want people to do try it out I guess try it out in the night list yeah so if you want to try it out just run towards.compile with Dynamic equals true on the in the torch compile call and so in the most recent night at least what you should should expect to work is inference will work and training will not work um Horus has a PR that makes trading work if you're feeling very adventurous you can go uh check that out but um it hasn't it hasn't landed to master yet we're going to land it before we actually go 2.0 and uh we are like passing like basically every Benchmark model on The Benchmark Suite if you're not actually compiling stuff that this is like kind of lame though because like the whole point of Partridge 2 is to compile stuff um we have a few more uh errors uh in compilation mostly around uh uh it's so silly um floor and seal don't work quite correctly if they show up in index computations like all time so we're hoping to fix that too before uh before brushcon but we'll see how that goes bleeding edge folks okay um all right so we haven't had any uh more questions so final call I guess Final Call uh questions um uh if we don't then we'll start uh wrapping up I I think we had a good discussion there's a lot of good insights shared and resources too so feel free to check those links out on the YouTube stream I'll also share them uh later on on on LinkedIn in addition to YouTube uh any more questions going once you know going what how do you do that never been an auction but uh going once going twice okay all right so I guess let's let's wrap up so uh thanks thanks Edward thanks um uh LS so we uh quick quick summary right we shared some resources and I'm gonna share them again and there's also a YouTube there's a video on YouTube on the conference talk from Horace you mentioned uh I'll be sharing all of those again many of those links are on the stream and uh please please check out Dev discuss and raise issues please try it out I think that's a single uh biggest call to action go try it out and uh let us know how it goes and uh we'll we'll close there then um see you in the next stream and last um sorry I'm going back on my file I just want to mention again the there are a lot more talks happening uh on all the individual features of pytos 2.0 they're on the pytus.org events page uh we've had several in the past and we are having we have more so you can go there and register or click the notify button on YouTube or subscribe whichever you prefer to be where to be notified and I'll close that all right um see everyone see you in the next live stream thanks Edwards thank you guys see you all later bye

Original Description

PyTorch 2.0 Q&A: 🗓️Feb 7th ⏰1pm PT 📝Dynamic Shapes & Calculating Maximum Batch Size Join Elias Ellison & Edward Yang to discuss symbolic shapes support in PyTorch. Follow the progress of Elias and Edward: https://dev-discuss.pytorch.org/t/state-of-symbolic-shapes-branch/777/35 Elias will also show use cases for this support (computing max batch sizes). Hosted by Developer Advocate Shashank Prasanna.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from PyTorch · PyTorch · 0 of 60

← Previous Next →
1 What is PyTorch?
What is PyTorch?
PyTorch
2 PyTorch Tutorial: A Quick Preview
PyTorch Tutorial: A Quick Preview
PyTorch
3 PyTorch Summer Hackathon 2019
PyTorch Summer Hackathon 2019
PyTorch
4 Tips and Tricks on Hacking with PyTorch: A Quick Tutorial by Brad Heintz
Tips and Tricks on Hacking with PyTorch: A Quick Tutorial by Brad Heintz
PyTorch
5 PyTorch 1.2 and PyTorch Hub: A Quick Introduction by Soumith Chintala and Ailing Zhang
PyTorch 1.2 and PyTorch Hub: A Quick Introduction by Soumith Chintala and Ailing Zhang
PyTorch
6 Torchtext 0.4 with Supervised Learning Datasets: A Quick Introduction by George Zhang
Torchtext 0.4 with Supervised Learning Datasets: A Quick Introduction by George Zhang
PyTorch
7 Torchaudio 0.3 with Kaldi Compatibility, New Transforms: A Quick Introduction by Jason Lian
Torchaudio 0.3 with Kaldi Compatibility, New Transforms: A Quick Introduction by Jason Lian
PyTorch
8 Torchvision 0.4 with Support for Video: A Quick Introduction by Francisco Massa
Torchvision 0.4 with Support for Video: A Quick Introduction by Francisco Massa
PyTorch
9 Introduction to Machine Learning for Developers at F8 2019
Introduction to Machine Learning for Developers at F8 2019
PyTorch
10 Powered by PyTorch at F8 2019
Powered by PyTorch at F8 2019
PyTorch
11 Developing and Scaling AI Experiences at Facebook with PyTorch at F8 2019
Developing and Scaling AI Experiences at Facebook with PyTorch at F8 2019
PyTorch
12 New Approaches to Image and Video Reconstruction Using Deep Learning at Facebook at F8 2019
New Approaches to Image and Video Reconstruction Using Deep Learning at Facebook at F8 2019
PyTorch
13 PyTorch Developer Conference 2018: Recap
PyTorch Developer Conference 2018: Recap
PyTorch
14 PyTorch Developer Conference 2018: Keynote & Deep Dive
PyTorch Developer Conference 2018: Keynote & Deep Dive
PyTorch
15 PyTorch Developer Conference 2018: Production & Research Sessions
PyTorch Developer Conference 2018: Production & Research Sessions
PyTorch
16 PyTorch Developer Conference 2018: Cloud & Academia Sessions
PyTorch Developer Conference 2018: Cloud & Academia Sessions
PyTorch
17 PyTorch Developer Conference 2018: Enterprise, Education, & Future of AI Panel
PyTorch Developer Conference 2018: Enterprise, Education, & Future of AI Panel
PyTorch
18 PyTorch Developer Conference 2019 | Full Livestream
PyTorch Developer Conference 2019 | Full Livestream
PyTorch
19 PyTorch Developer Conference 2019: Recap
PyTorch Developer Conference 2019: Recap
PyTorch
20 PyTorch Developer Conference Keynote - Mike Schroepfer
PyTorch Developer Conference Keynote - Mike Schroepfer
PyTorch
21 What’s new in PyTorch 1.3 - Lin Qiao
What’s new in PyTorch 1.3 - Lin Qiao
PyTorch
22 PyTorch Front-End Features: Named Tensors and Type Promotion - Gregory Chanan
PyTorch Front-End Features: Named Tensors and Type Promotion - Gregory Chanan
PyTorch
23 Research to Production: PyTorch JIT/TorchScript Updates - Michael Suo
Research to Production: PyTorch JIT/TorchScript Updates - Michael Suo
PyTorch
24 Quantization - Dmytro Dzhulgakov
Quantization - Dmytro Dzhulgakov
PyTorch
25 PyTorch ONNX Export Support - Lara Haidar, Microsoft
PyTorch ONNX Export Support - Lara Haidar, Microsoft
PyTorch
26 Apex -  Michael Carilli, NVIDIA
Apex - Michael Carilli, NVIDIA
PyTorch
27 Dataloader Design for PyTorch - Tongzhou Wang, MIT
Dataloader Design for PyTorch - Tongzhou Wang, MIT
PyTorch
28 Linear Algebra in PyTorch - Vishwak Srinivasan, CMU
Linear Algebra in PyTorch - Vishwak Srinivasan, CMU
PyTorch
29 PyTorch Mobile - David Reiss
PyTorch Mobile - David Reiss
PyTorch
30 Model Interpretability with Captum - Narine Kokhilkyan
Model Interpretability with Captum - Narine Kokhilkyan
PyTorch
31 Detectron2 - Next Gen Object Detection Library - Yuxin Wu
Detectron2 - Next Gen Object Detection Library - Yuxin Wu
PyTorch
32 Speech Extensions to Fairseq - Dmytro Okhonko
Speech Extensions to Fairseq - Dmytro Okhonko
PyTorch
33 PyTorch on Google Cloud TPUs - Google, Salesforce, Facebook
PyTorch on Google Cloud TPUs - Google, Salesforce, Facebook
PyTorch
34 PyTorch Summer Hackathon Winners - Joe Spisak, Sebastien Arnold, Tristan Deleu
PyTorch Summer Hackathon Winners - Joe Spisak, Sebastien Arnold, Tristan Deleu
PyTorch
35 PyTorch in Robotics - Yisong Yue, Caltech
PyTorch in Robotics - Yisong Yue, Caltech
PyTorch
36 StanfordNLP - Yuhao Zhang, Stanford
StanfordNLP - Yuhao Zhang, Stanford
PyTorch
37 Sotabench for Reproducible Research - Robert Stojnic, Papers with Code
Sotabench for Reproducible Research - Robert Stojnic, Papers with Code
PyTorch
38 Collaborative Natural Language Inference - Sasha Rush, Cornell
Collaborative Natural Language Inference - Sasha Rush, Cornell
PyTorch
39 Privacy Preserving AI - Andrew Trask, OpenMined
Privacy Preserving AI - Andrew Trask, OpenMined
PyTorch
40 CrypTen - Laurens van der Maaten
CrypTen - Laurens van der Maaten
PyTorch
41 PyTorch at Uber - Sidney Zhang, Uber
PyTorch at Uber - Sidney Zhang, Uber
PyTorch
42 PyTorch at Tesla - Andrej Karpathy, Tesla
PyTorch at Tesla - Andrej Karpathy, Tesla
PyTorch
43 PyTorch at Microsoft - Saurabh Tiwary, Microsoft
PyTorch at Microsoft - Saurabh Tiwary, Microsoft
PyTorch
44 PyTorch at Dolby Labs - Vivek Kumar, Dolby Labs
PyTorch at Dolby Labs - Vivek Kumar, Dolby Labs
PyTorch
45 PyTorch Developer Conference 2019 - Panel Discussion
PyTorch Developer Conference 2019 - Panel Discussion
PyTorch
46 Using deep learning and PyTorch to power next gen aircraft at Caltech
Using deep learning and PyTorch to power next gen aircraft at Caltech
PyTorch
47 Named Tensors, Model Quantization, and the Latest PyTorch Features - Part 1
Named Tensors, Model Quantization, and the Latest PyTorch Features - Part 1
PyTorch
48 TorchScript and PyTorch JIT | Deep Dive
TorchScript and PyTorch JIT | Deep Dive
PyTorch
49 Announcing the PyTorch Global Summer Hackathon 2020
Announcing the PyTorch Global Summer Hackathon 2020
PyTorch
50 Opening Up the Black Box: Model Understanding with Captum and PyTorch
Opening Up the Black Box: Model Understanding with Captum and PyTorch
PyTorch
51 PyTorch Mobile Runtime for Android
PyTorch Mobile Runtime for Android
PyTorch
52 Torchvision in 5 minutes
Torchvision in 5 minutes
PyTorch
53 3D Deep Learning with PyTorch3D
3D Deep Learning with PyTorch3D
PyTorch
54 What is Torchtext?
What is Torchtext?
PyTorch
55 TorchAudio: A Quick Intro
TorchAudio: A Quick Intro
PyTorch
56 PyTorch Mobile Runtime for iOS
PyTorch Mobile Runtime for iOS
PyTorch
57 PySlowFast: Deep learning with Video
PySlowFast: Deep learning with Video
PyTorch
58 PyTorch Pruning | How it's Made by Michela Paganini
PyTorch Pruning | How it's Made by Michela Paganini
PyTorch
59 Measuring Fairness in Machine Learning Systems
Measuring Fairness in Machine Learning Systems
PyTorch
60 PyTorch for Hackathons
PyTorch for Hackathons
PyTorch

Related Reads

Up next
RNNs Explained in 60 Seconds #ai #coding #machinelearning
Ascent
Watch →