PyTorch 2.0: TorchDynamo
Key Takeaways
PyTorch 2.0 introduces TorchDynamo, a graph capture technology that enables efficient compilation of PyTorch models, and discusses its importance for PyTorch 2.0's compilation story.
Full Transcript
do this [Music] I was waiting for all this talk for him to utter the word pie torch 2.0 and finally he said it welcome to pytorch conference this is my first one um I support the pi torch compiler team so as a compiler person PT 2.0 to a large extent is really a compeller story and you might wonder why does it take so long for us to actually fully Embrace compiler technology into the core of Pi torch and the reason is because what makes pie torch be loved by researchers is flexibility dynamism expressiveness is exactly what makes it very hard to compile so it took us five years of developing various Solutions compiler solutions to get to where we are and I want to call the inflection point of our journey searching for the right compiler solution polish Dynamo which is what I'm going to talk about so it starts with asking this question what is one thing that we could change first in the compiler stack that will fundamentally change the efficiency of Pi torch it's not a new compiler optimization it's not even a new compiler backend actually the first the key first move lies in the top part of the compiler which is the graph capture so we all know that graph mode execution can be a lot more efficient than executing Ops one at a time especially if you are a machine learning accelerator designer or if you are a production engineer unfortunately um for eager first machine learning framework like Pi torch it's actually non-trivial to get graphs reliably so that's why we care about graph capture because it's simple if you don't have reliable graph capture you don't have a reliable compiler coverage so let me introduce torch Dynamo our first out of box graph capture here's one example of a toy pytorch program and here is this one-liner change Torchlight compile and this in this example we actually allow you to specify a custom backend and in this particular example this my compiler basically takes the graph that's captured by Dynamo and do something about it which basically printed out and return a python callable so if you execute this program you'll be surprised you see actually three graphs being printed out because in this example we actually put a Twist of a data dependent control flow so this graph break actually is the feature it's a fundamental design feature of torch Dynamo that makes it out of box in sound so let me double click on what makes torch Dynamo sound and outer box number one is partial graph capture so this is our ability to skip anything unwanted and then fall back to eager and with partial graph capture we save the programmers to have to shoehorsing their models into what the back-end compiler want instead we'll just seem to be translating out in and out of between eager and the graph mode and then there is the guarded graph generation basically we want to make sure whatever we're capturing at capture time is still valid when we're replaying this graph and finally if we did find a mismatch on the guard we have the ability to just in time recapture them so basically partial guarded graph capture with just in-time recapturing is what set Dynamo apart from all our previous graph capture Technologies and by the way there's this thing called aot autograph that captures the backward part of the graph and that makes it working for training so I want to show you some numbers which assume is already highlighted in his keynote that would back the claims we made so far you have seen the 7K plus model so this demonstrates that it indeed works out of box and right now we have more than 20 experimental backends for inference that's integrated with torch Dynamo and we have more than one backhand that's integrated with torchamel for training and the number is smaller on trading side now because of Dynamo but because they're just less number of training back-ends available and we're working actively with vendors to migrate into this new stack and finally you've seen this 30 number what does it say actually it's not only talking about the performance it also demonstrates that partial graph actually works so by capturing crafts and allowing graph breaks we not only did not diminish the capability of the back end optimizing compiler we actually amplify it by bringing more models or more regions of the models into the compiler world so Dynamo is great but if you are an experienced pytorch compiler user you would ask this question so which technology do I use as of today we actually recommend using torture animal for any scenarios that can be deployed with eager so that includes most of the training scenario and some of the inference scenarios and for all the other scenarios which we call the exports past scenario our existing solution still works but in the future actually not so far future we actually want to have a con organic consolidation uh of front-end technology around torch Dynamo um it's it has always been our vision to have a unifying front end and have a vibrant thriving back-end ecosystem and the speed of our consolidation actually depends on two things number one how fast do we deliver the pytorch 2.0 export story and number two is how fast the vendors can migrate into this new stack so we look forward to that new experience of a Consolidated front end yeah so that's Dynamo and I want to bring you back to the question we asked in the beginning of the talk what is that one thing right so as a compiler optimizing compiler person for two decades and never imagine myself talk about a computer friend end but the front end part of the compiler graph capture really matters because once we're able to bring most of the pytorch models simply into the graph world the era of compiler accelerated pytorch really arrives so if you are excited about this Vision there are more Nuance aspects of the compiler story that you are going to hear today and in the following months and if you're a ml practitioner I would encourage you to try out pt2 on your workload and if you're a vendor I encourage you to talk to us to explore migrating into this new stack and finally we have some talk series that would dive deeper into the technology in in the pytorch 2.0 stack so with that I'm gonna pass the torch to Jason Enzo [Applause] thank you
Original Description
Peng Wu speaks at PyTorch Conference about TorchDynamo, the first out-of-the-box Graph Capture for PyTorch and why it matters.
This short talk shares highlights on TorchDynamo, PT 2.0's new graph capture technology. It answers questions like:
1) Why is graph capture so important for PT2's compiler stack?
2) How is TorchDynamo better than previous PyTorch graph captures?
3) What is the current state of TorchDynamo?
3) What are our short- and long-term recommendations for vendors/users?
Visit our website: https://pytorch.org/
Read our blog: https://pytorch.org/blog/
Follow us on Twitter: https://twitter.com/PyTorch
Follow us on LinkedIn: https://www.linkedin.com/company/pyto...
Follow us on Facebook: https://www.facebook.com/pytorch
#PyTorch #ArtificialIntelligence #MachineLearning
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from PyTorch · PyTorch · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
What is PyTorch?
PyTorch
PyTorch Tutorial: A Quick Preview
PyTorch
PyTorch Summer Hackathon 2019
PyTorch
Tips and Tricks on Hacking with PyTorch: A Quick Tutorial by Brad Heintz
PyTorch
PyTorch 1.2 and PyTorch Hub: A Quick Introduction by Soumith Chintala and Ailing Zhang
PyTorch
Torchtext 0.4 with Supervised Learning Datasets: A Quick Introduction by George Zhang
PyTorch
Torchaudio 0.3 with Kaldi Compatibility, New Transforms: A Quick Introduction by Jason Lian
PyTorch
Torchvision 0.4 with Support for Video: A Quick Introduction by Francisco Massa
PyTorch
Introduction to Machine Learning for Developers at F8 2019
PyTorch
Powered by PyTorch at F8 2019
PyTorch
Developing and Scaling AI Experiences at Facebook with PyTorch at F8 2019
PyTorch
New Approaches to Image and Video Reconstruction Using Deep Learning at Facebook at F8 2019
PyTorch
PyTorch Developer Conference 2018: Recap
PyTorch
PyTorch Developer Conference 2018: Keynote & Deep Dive
PyTorch
PyTorch Developer Conference 2018: Production & Research Sessions
PyTorch
PyTorch Developer Conference 2018: Cloud & Academia Sessions
PyTorch
PyTorch Developer Conference 2018: Enterprise, Education, & Future of AI Panel
PyTorch
PyTorch Developer Conference 2019 | Full Livestream
PyTorch
PyTorch Developer Conference 2019: Recap
PyTorch
PyTorch Developer Conference Keynote - Mike Schroepfer
PyTorch
What’s new in PyTorch 1.3 - Lin Qiao
PyTorch
PyTorch Front-End Features: Named Tensors and Type Promotion - Gregory Chanan
PyTorch
Research to Production: PyTorch JIT/TorchScript Updates - Michael Suo
PyTorch
Quantization - Dmytro Dzhulgakov
PyTorch
PyTorch ONNX Export Support - Lara Haidar, Microsoft
PyTorch
Apex - Michael Carilli, NVIDIA
PyTorch
Dataloader Design for PyTorch - Tongzhou Wang, MIT
PyTorch
Linear Algebra in PyTorch - Vishwak Srinivasan, CMU
PyTorch
PyTorch Mobile - David Reiss
PyTorch
Model Interpretability with Captum - Narine Kokhilkyan
PyTorch
Detectron2 - Next Gen Object Detection Library - Yuxin Wu
PyTorch
Speech Extensions to Fairseq - Dmytro Okhonko
PyTorch
PyTorch on Google Cloud TPUs - Google, Salesforce, Facebook
PyTorch
PyTorch Summer Hackathon Winners - Joe Spisak, Sebastien Arnold, Tristan Deleu
PyTorch
PyTorch in Robotics - Yisong Yue, Caltech
PyTorch
StanfordNLP - Yuhao Zhang, Stanford
PyTorch
Sotabench for Reproducible Research - Robert Stojnic, Papers with Code
PyTorch
Collaborative Natural Language Inference - Sasha Rush, Cornell
PyTorch
Privacy Preserving AI - Andrew Trask, OpenMined
PyTorch
CrypTen - Laurens van der Maaten
PyTorch
PyTorch at Uber - Sidney Zhang, Uber
PyTorch
PyTorch at Tesla - Andrej Karpathy, Tesla
PyTorch
PyTorch at Microsoft - Saurabh Tiwary, Microsoft
PyTorch
PyTorch at Dolby Labs - Vivek Kumar, Dolby Labs
PyTorch
PyTorch Developer Conference 2019 - Panel Discussion
PyTorch
Using deep learning and PyTorch to power next gen aircraft at Caltech
PyTorch
Named Tensors, Model Quantization, and the Latest PyTorch Features - Part 1
PyTorch
TorchScript and PyTorch JIT | Deep Dive
PyTorch
Announcing the PyTorch Global Summer Hackathon 2020
PyTorch
Opening Up the Black Box: Model Understanding with Captum and PyTorch
PyTorch
PyTorch Mobile Runtime for Android
PyTorch
Torchvision in 5 minutes
PyTorch
3D Deep Learning with PyTorch3D
PyTorch
What is Torchtext?
PyTorch
TorchAudio: A Quick Intro
PyTorch
PyTorch Mobile Runtime for iOS
PyTorch
PySlowFast: Deep learning with Video
PyTorch
PyTorch Pruning | How it's Made by Michela Paganini
PyTorch
Measuring Fairness in Machine Learning Systems
PyTorch
PyTorch for Hackathons
PyTorch
More on: ML Pipelines
View skill →Related Reads
📰
📰
📰
📰
Help Choosing Neural Network Architecture for Matrix Classification
Reddit r/deeplearning
How to Choose the Best Deep Learning Model for Medical Imaging
Medium · Deep Learning
Another Way to Read Neural Geometry
Medium · Data Science
Another Way to Read Neural Geometry
Medium · Deep Learning
🎓
Tutor Explanation
DeepCamp AI