PyTorch 2.0: TorchInductor

PyTorch · Intermediate ·📰 AI News & Updates ·3y ago

Key Takeaways

The video discusses TorchInductor, a PyTorch-native compiler backend for PyTorch 2.0, and its key principles, technologies, and results.

Full Transcript

[Music] [Music] thank you hello everyone my name is Jason Ansel and I'm going to talk about torch inductor which is a new compiler back end for pi torch 2.0 when designing torch inductor we started off with three key principles so the first principle is that torch inductor is pi torch native and what this means is that inductor uses very similar abstractions to uh Pi torch eager which allows it to Faithfully capture all the behavior of pytorch the next key principle is python first so with pytorch 2 we are embracing Python and part of that involves writing inductor entirely in Python which makes it much easier to hack on and extend and then finally uh we wanted inductor to be really General and so we focused on breadth rather than depth early on which means tackling the tricky operators and tricky optimizations early to ensure that our design was General and could scale inductor uses uh three key Technologies as well the first technology is a defined by run Loop level AR and what that means is that the core compiler IR that's done at the Loop level is actually a python callable and to do things like code generation and Analysis we actually execute the IR and do that which is a sort of Novel trick in the compiler space but some things that's been used um for pytorch programs uh for a while uh next we we really wanted inductor to support Dynamic shapes and strides um from uh day one and so inductor uses Senpai which is a symbolic math library to reason about uh shapes and generate code that's not uh specialized uh to specific input sizes and finally I'm going to reinvent the wheel we wanted to reuse uh state-of-the-art languages and so we took inspiration from what our users were doing and increasingly we were seeing people write high performance kernels in this new language Triton um by uh Philip tillett at open Ai and uh so what we have is we have a compiler in torch inductor which generates uh Triton code which is easy to understand you can look at the output and and inspect what it does or even change it and on CPUs we generate C plus plus code uh so what has Triton uh Triton is you could think of it as a better Cuda so uh it's a higher level language than Cuda but also lower level than a lot of pre-existing dsls and it makes it really easy to write high performance code even for tricky things like Matrix multiplies that are competitive uh with things like kudi and n and kublas next I'm going to talk about Mine by run Loop level AR and so here's what the IR for a permute and an ad would look like and so we have this inner function here which takes a list of senpai Expressions that list of symbi Expressions represents a coordinate um uh symbolically that we want to generate and then the in the the body of this function we'll call ops.load twice and ops.add once and the way you use this IR is for example if you wanted to write an analysis pass you could replace This Global variable Ops with something which which records the loads then you could run this function and you could uh do your analysis to figure out uh what loads this function makes and if you wanted to do a code generation you could replace this Ops function with something that printed out Trident code execute this this function and do code gen and this makes it really easy to write lowerings because if you're building composite Ops that sort of take two inputs together you can actually use the features of the Python language to build your lowering and then part way through the compilation process we actually use FX to trace this defined by run IR which gives us a IR that's that's more manipulable so here's an overview of the compiler stack uh so uh we start off with the graph that George Dynamo has captured and we use aot autograd and Prim torch to decompose the operator set into a much more minimal operator set around 250 operators and aot autograd captures the forwards and the backwards graph which we compile uh independently next there's uh graph lowerings graph lowerings takes this primitive operator set which is around 250 operators and lowers it to inductors Loop level ir and the inductors Looper level contains only around 50 operators so we've removed a lot of the complexity and simplified the graph substantially uh and uh we're now operating on how to on basically single elements of tensors rather than full tensors uh as in the input next there's the scheduling phase scheduling is where we decide what gets fused with what and do other optimizations such as memory planning and tiling and then finally we have code generation and so there's two pieces to code generation we have the backend code generation which either generates Trident code or C plus plus and then we have what we call wrapper code gen and wrapper Cogen is the code that stitches together the calls to many different kernels and it basically replaces The Interpreter part of the compiler stack so here's some results on GPU I know Sue Miss shared these earlier but inductor generates up to 1.86 Geo mean speed up on large realistic Benchmark Suites and we're super excited about these results and hope to continue to improve them in the future we also have results on on CPU and these CPU results are a collaboration with the Intel Pi torch team we see up to 1.26 X geomine speed up on uh on on CPU inference and this serves the purpose of uh having inductor support CPUs but it also supports the purpose of making sure the inductor is General we didn't want to build a compiler that only supported gpus we wanted something that could scale to support a wide variety of Hardware back ends and having a c plus as well as it is Trident forces that generality so thanks so much uh if you uh the the code base for inductor is in the pi torch repo I've linked it here and there's more information uh in the blog post as well thanks so much [Music]

Original Description

Jason Ansel speaks at PyTorch Conference 2022 about TorchInductor, a PyTorch-native compiler. TorchInductor is a deep learning compiler backend for PyTorch 2.0. For NVIDIA GPUs, it uses OpenAI Triton as a key building block. TorchInductor uses similar abstractions to PyTorch eager, and is general purpose enough to support the wide breadth of features in PyTorch. TorchInductor uses a pythonic define-by-run loop level IR to automatically map PyTorch models into generated Triton code on GPUs and C++/OpenMP on CPUs. TorchInductor’s core loop level IR contains only ~50 operators, and it is implemented in Python, making it easily hackable and extensible. Visit our website: https://pytorch.org/ Read our blog: https://pytorch.org/blog/ Follow us on Twitter: https://twitter.com/PyTorch Follow us on LinkedIn: https://www.linkedin.com/company/pyto... Follow us on Facebook: https://www.facebook.com/pytorch #PyTorch #ArtificialIntelligence #MachineLearning
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from PyTorch · PyTorch · 0 of 60

← Previous Next →
1 What is PyTorch?
What is PyTorch?
PyTorch
2 PyTorch Tutorial: A Quick Preview
PyTorch Tutorial: A Quick Preview
PyTorch
3 PyTorch Summer Hackathon 2019
PyTorch Summer Hackathon 2019
PyTorch
4 Tips and Tricks on Hacking with PyTorch: A Quick Tutorial by Brad Heintz
Tips and Tricks on Hacking with PyTorch: A Quick Tutorial by Brad Heintz
PyTorch
5 PyTorch 1.2 and PyTorch Hub: A Quick Introduction by Soumith Chintala and Ailing Zhang
PyTorch 1.2 and PyTorch Hub: A Quick Introduction by Soumith Chintala and Ailing Zhang
PyTorch
6 Torchtext 0.4 with Supervised Learning Datasets: A Quick Introduction by George Zhang
Torchtext 0.4 with Supervised Learning Datasets: A Quick Introduction by George Zhang
PyTorch
7 Torchaudio 0.3 with Kaldi Compatibility, New Transforms: A Quick Introduction by Jason Lian
Torchaudio 0.3 with Kaldi Compatibility, New Transforms: A Quick Introduction by Jason Lian
PyTorch
8 Torchvision 0.4 with Support for Video: A Quick Introduction by Francisco Massa
Torchvision 0.4 with Support for Video: A Quick Introduction by Francisco Massa
PyTorch
9 Introduction to Machine Learning for Developers at F8 2019
Introduction to Machine Learning for Developers at F8 2019
PyTorch
10 Powered by PyTorch at F8 2019
Powered by PyTorch at F8 2019
PyTorch
11 Developing and Scaling AI Experiences at Facebook with PyTorch at F8 2019
Developing and Scaling AI Experiences at Facebook with PyTorch at F8 2019
PyTorch
12 New Approaches to Image and Video Reconstruction Using Deep Learning at Facebook at F8 2019
New Approaches to Image and Video Reconstruction Using Deep Learning at Facebook at F8 2019
PyTorch
13 PyTorch Developer Conference 2018: Recap
PyTorch Developer Conference 2018: Recap
PyTorch
14 PyTorch Developer Conference 2018: Keynote & Deep Dive
PyTorch Developer Conference 2018: Keynote & Deep Dive
PyTorch
15 PyTorch Developer Conference 2018: Production & Research Sessions
PyTorch Developer Conference 2018: Production & Research Sessions
PyTorch
16 PyTorch Developer Conference 2018: Cloud & Academia Sessions
PyTorch Developer Conference 2018: Cloud & Academia Sessions
PyTorch
17 PyTorch Developer Conference 2018: Enterprise, Education, & Future of AI Panel
PyTorch Developer Conference 2018: Enterprise, Education, & Future of AI Panel
PyTorch
18 PyTorch Developer Conference 2019 | Full Livestream
PyTorch Developer Conference 2019 | Full Livestream
PyTorch
19 PyTorch Developer Conference 2019: Recap
PyTorch Developer Conference 2019: Recap
PyTorch
20 PyTorch Developer Conference Keynote - Mike Schroepfer
PyTorch Developer Conference Keynote - Mike Schroepfer
PyTorch
21 What’s new in PyTorch 1.3 - Lin Qiao
What’s new in PyTorch 1.3 - Lin Qiao
PyTorch
22 PyTorch Front-End Features: Named Tensors and Type Promotion - Gregory Chanan
PyTorch Front-End Features: Named Tensors and Type Promotion - Gregory Chanan
PyTorch
23 Research to Production: PyTorch JIT/TorchScript Updates - Michael Suo
Research to Production: PyTorch JIT/TorchScript Updates - Michael Suo
PyTorch
24 Quantization - Dmytro Dzhulgakov
Quantization - Dmytro Dzhulgakov
PyTorch
25 PyTorch ONNX Export Support - Lara Haidar, Microsoft
PyTorch ONNX Export Support - Lara Haidar, Microsoft
PyTorch
26 Apex -  Michael Carilli, NVIDIA
Apex - Michael Carilli, NVIDIA
PyTorch
27 Dataloader Design for PyTorch - Tongzhou Wang, MIT
Dataloader Design for PyTorch - Tongzhou Wang, MIT
PyTorch
28 Linear Algebra in PyTorch - Vishwak Srinivasan, CMU
Linear Algebra in PyTorch - Vishwak Srinivasan, CMU
PyTorch
29 PyTorch Mobile - David Reiss
PyTorch Mobile - David Reiss
PyTorch
30 Model Interpretability with Captum - Narine Kokhilkyan
Model Interpretability with Captum - Narine Kokhilkyan
PyTorch
31 Detectron2 - Next Gen Object Detection Library - Yuxin Wu
Detectron2 - Next Gen Object Detection Library - Yuxin Wu
PyTorch
32 Speech Extensions to Fairseq - Dmytro Okhonko
Speech Extensions to Fairseq - Dmytro Okhonko
PyTorch
33 PyTorch on Google Cloud TPUs - Google, Salesforce, Facebook
PyTorch on Google Cloud TPUs - Google, Salesforce, Facebook
PyTorch
34 PyTorch Summer Hackathon Winners - Joe Spisak, Sebastien Arnold, Tristan Deleu
PyTorch Summer Hackathon Winners - Joe Spisak, Sebastien Arnold, Tristan Deleu
PyTorch
35 PyTorch in Robotics - Yisong Yue, Caltech
PyTorch in Robotics - Yisong Yue, Caltech
PyTorch
36 StanfordNLP - Yuhao Zhang, Stanford
StanfordNLP - Yuhao Zhang, Stanford
PyTorch
37 Sotabench for Reproducible Research - Robert Stojnic, Papers with Code
Sotabench for Reproducible Research - Robert Stojnic, Papers with Code
PyTorch
38 Collaborative Natural Language Inference - Sasha Rush, Cornell
Collaborative Natural Language Inference - Sasha Rush, Cornell
PyTorch
39 Privacy Preserving AI - Andrew Trask, OpenMined
Privacy Preserving AI - Andrew Trask, OpenMined
PyTorch
40 CrypTen - Laurens van der Maaten
CrypTen - Laurens van der Maaten
PyTorch
41 PyTorch at Uber - Sidney Zhang, Uber
PyTorch at Uber - Sidney Zhang, Uber
PyTorch
42 PyTorch at Tesla - Andrej Karpathy, Tesla
PyTorch at Tesla - Andrej Karpathy, Tesla
PyTorch
43 PyTorch at Microsoft - Saurabh Tiwary, Microsoft
PyTorch at Microsoft - Saurabh Tiwary, Microsoft
PyTorch
44 PyTorch at Dolby Labs - Vivek Kumar, Dolby Labs
PyTorch at Dolby Labs - Vivek Kumar, Dolby Labs
PyTorch
45 PyTorch Developer Conference 2019 - Panel Discussion
PyTorch Developer Conference 2019 - Panel Discussion
PyTorch
46 Using deep learning and PyTorch to power next gen aircraft at Caltech
Using deep learning and PyTorch to power next gen aircraft at Caltech
PyTorch
47 Named Tensors, Model Quantization, and the Latest PyTorch Features - Part 1
Named Tensors, Model Quantization, and the Latest PyTorch Features - Part 1
PyTorch
48 TorchScript and PyTorch JIT | Deep Dive
TorchScript and PyTorch JIT | Deep Dive
PyTorch
49 Announcing the PyTorch Global Summer Hackathon 2020
Announcing the PyTorch Global Summer Hackathon 2020
PyTorch
50 Opening Up the Black Box: Model Understanding with Captum and PyTorch
Opening Up the Black Box: Model Understanding with Captum and PyTorch
PyTorch
51 PyTorch Mobile Runtime for Android
PyTorch Mobile Runtime for Android
PyTorch
52 Torchvision in 5 minutes
Torchvision in 5 minutes
PyTorch
53 3D Deep Learning with PyTorch3D
3D Deep Learning with PyTorch3D
PyTorch
54 What is Torchtext?
What is Torchtext?
PyTorch
55 TorchAudio: A Quick Intro
TorchAudio: A Quick Intro
PyTorch
56 PyTorch Mobile Runtime for iOS
PyTorch Mobile Runtime for iOS
PyTorch
57 PySlowFast: Deep learning with Video
PySlowFast: Deep learning with Video
PyTorch
58 PyTorch Pruning | How it's Made by Michela Paganini
PyTorch Pruning | How it's Made by Michela Paganini
PyTorch
59 Measuring Fairness in Machine Learning Systems
Measuring Fairness in Machine Learning Systems
PyTorch
60 PyTorch for Hackathons
PyTorch for Hackathons
PyTorch

TorchInductor is a PyTorch-native compiler backend for PyTorch 2.0 that uses OpenAI Triton and Senpai to generate high-performance code. It is designed to be general, Python-first, and supports dynamic shapes and strides.

Key Takeaways
  1. Design a compiler with PyTorch-native abstractions
  2. Use Defined-by-run Loop level IR for code generation and analysis
  3. Implement Senpai for symbolic math and dynamic shape support
  4. Generate Triton code for NVIDIA GPUs and C++ code for CPUs
  5. Optimize performance with scheduling and memory planning
💡 TorchInductor's use of Defined-by-run Loop level IR and Senpai enables efficient code generation and analysis for PyTorch models.

Related Reads

📰
Musk thanks Micron for chips, and builds a $55bn fab to replace it
Elon Musk thanks Micron for providing Tesla with a significant allocation of memory chips, highlighting the scarcity of this crucial component in the AI boom
The Next Web AI
📰
Jensen Huang calls the AI jobs panic ‘complete nonsense’, and takes aim at his peers
Nvidia CEO Jensen Huang dismisses AI job replacement panic as 'complete nonsense', offering a contrasting view to his peers
The Next Web AI
📰
IMF says Africa has to keep lights on before it can bet on AI
Africa's AI ambitions are hindered by unreliable electricity, highlighting the need for basic infrastructure before investing in AI
TechCabal
📰
What Does Job Security Even Look Like In 2026? It Starts With Skills
Job security in 2026 requires adapting to AI and economic uncertainty by acquiring in-demand skills
Forbes Innovation
Up next
Claude's Small Business Plugin: 31 Workflows You Can Run Today
Kevin Farugia AI Automation
Watch →