[TVM] Universally Deploy Large-language Models via ML Compilation - Tianqi Chen, CMU & OctoAI

PyTorch · Beginner ·📐 ML Fundamentals ·1y ago
Skills: ML Pipelines80%

Key Takeaways

Deploys large-language models via ML compilation using TVM

Full Transcript

okay let's get started so hello everyone my name is teni CH so I today I'm going to tell you a bit about our recent effort in apach TVM and this General effort on deploying large language models which Universal compilation so it's great to be here and talk about machine learning compilations and we've been working on M Mo compiler for several years now and it's great to say compiler starting getting being used in a lot of places but one of the things to recap and we think about this field is what are the biggest challenges and I think of earlier speakers also alluded to that in a sense that when we are doing machine learning development and machine learning engineering the big challenge is not only trying to build a solution for today but continuously to keep up with you know hey I have a new model I have a new demand new customization coming in and you want to be able to keep up with that and that's the biggest challenge in here today we going to talk about three high level thoughts that are related to machine learning Computing constructions um and uh and you know we're also going to talk about the specific domain specific compiler we are building in the past one year first first point is that we want to think about machine learning compilation research and development as of now are being can be quite different from a traditional compiler development in a sense that you traditionally if you thinking about compiler you developed it for several years you ship it right and then you make use of it for for know for for effectively quite a while um and if you are a machine Lear scientist what you do is you go and developed a material framework on top of uh uh develop your models on top of material Foundation you you don't care about the blackbox it's going to go and work for you however I would argue that machine Lear competition development itself should be more productive in a sense that you know why don't we go and build a language model compiler in a month and you know if you can do that that allows you to be able to go and iterate together with the model development in some sense and uh with a lot of the new emerging workload today actually we are getting into uh we are see in need of you know hey we can build a specific solution for a particular model or particular kind of model right and the Really bottom neck there is actually the machine learning compilation infrastructure need to learn a lot from actually machine learning Frameworks where you know with py to you will be able to Define your new model and R it in a day right and ideally you know we should be able to start building some of R&D toward that direction so uh one of the lessons we learned so far is we would like to actually make python as a first class citizen in ml compiler development so idea is that you know want to be able to embed most of a unified IR within a python python you know EST and you will be able to go and manipulate a lot of that Additionally you know you want to build modular composition transformations in Python if necessary and you know if you find them being useful you put them into C++ or more material language and and make it more efficient right and in that sense our development looks more like this you have a existing pipeline that works well for hopefully for most of use cases you start to look at a new model maybe okay hey I need a new attention operator right previously you need to go to your C+ code base you dig into okay here's the place I lower attention I go and Hack That code and then you maybe after month you shi it and then you observe a performance again or maybe you know you get a feedback or this is not really want right so instead what we really want is you know you you introduce start to introduce python based pipeline libraries you say dis patch to my attention uh and the customize your pipeline and then you compose it with the rest of a pipeline ideally and it gives you immediate feedback maybe that particular pass is only working for one particular model even but you know if it take one day you customiz pipeline you ship it if it works that's great if not you know let's come back and re reiterate so so that's a high level uh in terms of we really want to be able to make M compilation development more productive and uh you know making it as part of our machine learning engineering pipeline okay the second thing is about some of high level lessons we learn on building the media representation specifically um along the cross level obstruction and first class dynamation so first of all a lot of our talks they also mentioned about different levels of of you know compil obstructions for example if you want to look at a kernel language usually kernel language at what we call Loop level or tensor program level that describes high level Loops Loop nestings you know and GPU thre bindings and so on right and we also talk about graph compilers you says you have graph competitional graph node where each of a node contains high level tens of operations and so on usually what we do is we start with graph level compiler you optimize it in graph level and you lower into the kernel level and you go and optimize them what we really find is actually helpful to try to bring them together in a more organic way and specifically you want to be able to both analyze the the loop level program to automatically provides you annotation which otherwise you have to manually annotate for each of the compination graph not as well as do some class level transformation so I will because of time I'm going to show one particular case study which is you know if you about language models you want to look at different kind of quation format like awq Q4 fp8 fp4 and the first problem you will find is that you know if you started work with a compiler it does not support that particular quation format it's very common like I think you know I I face that when we start our old compiler pipeline so if you have a Closs level representation what you can do is actually we can use a lower level tensor program to directly represent hey here's how I'm going to decode a Q4 into my floating point and then ideally if you you can create a cross- level uh optimization that you know automatically fusees the low level tenser code with the original maal code and do certain manipulation so on so having Closs level really brings some kind of flexibility of course this still is possible if you go and change your compiler introduce a new data data type introduce operators introduce lowering rules and uh and so on but know with this cross level thing we'll be able to add that in in in one day okay the second thing we find is really helpful is we really want to bring first class symbolic shape because language model you have Dynamic shape uh both in terms of B size as well as you know sequence lens and so on right and there are two different kind of ways normally you handle this so first way is you try to do what we call any ship so you have a dyamic ship Dimension you mark as a question mark I don't know what it is symbolic ship means that you have a you have a symbolic Dimension and after I flatten it right is n by4 after flatten it become n * 4 so compiler will be able to reason about their relations like I will be able to tell that the number of elements here is going to be same number element in the second part right so as of a result maybe in memory allocations I will be able to tie them together because they have the same number of elements so bringing first class symbolic ship is really useful but it's also really challenging specifically in our case we really want to not only target just in time compilation but also ahead of time compilation so we want to be able to Target full program right so that means that we want to be able to enable global tracking including having function signatures that contains like high level relations like um taking a tensor by n by4 and after calling gives you tensor by n * 4 and additionally you want to be able to propagate this symbolic shape relations both in terms of the competition graph level as well as your low level tens level right so that uh gives you better information okay so with all that element together I want to talk about the specific use case we have been building the use cases we really want to be able to enable deploying large language model universally across different environments and to be clear this is not a general purpose compiler it's this a spef civico compiler developed for Lang language models built on top of the uh uh compiler infrastructure I talk about in the first part and mm is uh you know as of now it supports all the hardare Targets in here uh you can run on Windows Linux Mac across platforms uh it have a open a compatible server that gives you know allows you to run different models um it runs on iOS if you go to App Store search for M chat you'll be able to find a demo app that can run llama models and and other small Lang models directly on iPhone uh it runs on Android you can bring it to you know $100 orange pie device it runs on Steam deck uh if you want to play games and additionally we're also building a unified engine for doing additional things like efficient structure generation where it have building support for Json grammar and gives you new zero overhead generation here and one of the goal again is we really want to be able to unify the overall compilation and generation for both Cloud cases and edge cases so so far I talk about a lot of edge cases more recently we have been also working on some of our low latency server optimizations so in this particular case it's a low latency server Benchmark uh checking you know in this case we are talking about both in terms of throughput highest better and this output token part and a lot of Benchmark recently in server focus on stut which means that you know you want to have push the batch size to be as large as possible so that you can you know have utilize the GPU as much as possible however most of the GPU providers most of language model providers also care about about you know output tokens right you really want to give your users as fast response as possible so really what you want to look at example if you want to have your output token uh to be above 100 100 token per second you want to look at something like everything about uh on the left left hand side of this lines it's useful and you can find that you know with enough optimization compilation so you can get pretty decent performance as as good as you know what's the best out there on a server cases and you can also make it run on most of the environments uh one of my favorite example you can go and you know directly run our web browsers actually I'm going to show a live demo now uh with last one minute uh so if you go to chat. web. you can directly say you know tell me about pyos given this is a pyos conference and what's going to do is it's always a danger to show live demos let me Let me refresh uh try again and and you can try it on a laptop as well uh and what I try to do is we're going to fetch the weights on to a web browser cach it runs all locally on my uh on my laptop and it's it's being populated by islama model by the way thanks Facebook uh and you know it's as uh compiles actually the mm compiles it to web assembly and web GPU so you can you can say run pretty fast and you know actually you don't you just need to open a web page to to run okay so that's uh and coming back to presentation so yeah so or um all the materials um in today St you can find in this website called m.ai I'm also really passionate about you know actually making machine learning compilation more accessible and learn more about it so I'm creating a course about M Mo compilation in general and uh the uh overall specific uh domain specific compiler mm is going to be available on that website okay hopefully I'm just in time

Original Description

[TVM] Universally Deploy Large-language Models via ML Compilation - Tianqi Chen, CMU & OctoAI Deploying deep learning models on various devices has become an important topic. Machine learning compilation is an emerging field that leverages compiler and automatic search techniques to accelerate AI models. ML compilation brings a unique set of challenges: emerging machine learning models; increasing hardware specialization brings a diverse set of acceleration primitives; growing tension between flexibility and performance. In this talk. I then discuss our experience in bringing foundational models to a variety of devices and hardware environments through machine learning compilation.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from PyTorch · PyTorch · 0 of 60

← Previous Next →
1 What is PyTorch?
What is PyTorch?
PyTorch
2 PyTorch Tutorial: A Quick Preview
PyTorch Tutorial: A Quick Preview
PyTorch
3 PyTorch Summer Hackathon 2019
PyTorch Summer Hackathon 2019
PyTorch
4 Tips and Tricks on Hacking with PyTorch: A Quick Tutorial by Brad Heintz
Tips and Tricks on Hacking with PyTorch: A Quick Tutorial by Brad Heintz
PyTorch
5 PyTorch 1.2 and PyTorch Hub: A Quick Introduction by Soumith Chintala and Ailing Zhang
PyTorch 1.2 and PyTorch Hub: A Quick Introduction by Soumith Chintala and Ailing Zhang
PyTorch
6 Torchtext 0.4 with Supervised Learning Datasets: A Quick Introduction by George Zhang
Torchtext 0.4 with Supervised Learning Datasets: A Quick Introduction by George Zhang
PyTorch
7 Torchaudio 0.3 with Kaldi Compatibility, New Transforms: A Quick Introduction by Jason Lian
Torchaudio 0.3 with Kaldi Compatibility, New Transforms: A Quick Introduction by Jason Lian
PyTorch
8 Torchvision 0.4 with Support for Video: A Quick Introduction by Francisco Massa
Torchvision 0.4 with Support for Video: A Quick Introduction by Francisco Massa
PyTorch
9 Introduction to Machine Learning for Developers at F8 2019
Introduction to Machine Learning for Developers at F8 2019
PyTorch
10 Powered by PyTorch at F8 2019
Powered by PyTorch at F8 2019
PyTorch
11 Developing and Scaling AI Experiences at Facebook with PyTorch at F8 2019
Developing and Scaling AI Experiences at Facebook with PyTorch at F8 2019
PyTorch
12 New Approaches to Image and Video Reconstruction Using Deep Learning at Facebook at F8 2019
New Approaches to Image and Video Reconstruction Using Deep Learning at Facebook at F8 2019
PyTorch
13 PyTorch Developer Conference 2018: Recap
PyTorch Developer Conference 2018: Recap
PyTorch
14 PyTorch Developer Conference 2018: Keynote & Deep Dive
PyTorch Developer Conference 2018: Keynote & Deep Dive
PyTorch
15 PyTorch Developer Conference 2018: Production & Research Sessions
PyTorch Developer Conference 2018: Production & Research Sessions
PyTorch
16 PyTorch Developer Conference 2018: Cloud & Academia Sessions
PyTorch Developer Conference 2018: Cloud & Academia Sessions
PyTorch
17 PyTorch Developer Conference 2018: Enterprise, Education, & Future of AI Panel
PyTorch Developer Conference 2018: Enterprise, Education, & Future of AI Panel
PyTorch
18 PyTorch Developer Conference 2019 | Full Livestream
PyTorch Developer Conference 2019 | Full Livestream
PyTorch
19 PyTorch Developer Conference 2019: Recap
PyTorch Developer Conference 2019: Recap
PyTorch
20 PyTorch Developer Conference Keynote - Mike Schroepfer
PyTorch Developer Conference Keynote - Mike Schroepfer
PyTorch
21 What’s new in PyTorch 1.3 - Lin Qiao
What’s new in PyTorch 1.3 - Lin Qiao
PyTorch
22 PyTorch Front-End Features: Named Tensors and Type Promotion - Gregory Chanan
PyTorch Front-End Features: Named Tensors and Type Promotion - Gregory Chanan
PyTorch
23 Research to Production: PyTorch JIT/TorchScript Updates - Michael Suo
Research to Production: PyTorch JIT/TorchScript Updates - Michael Suo
PyTorch
24 Quantization - Dmytro Dzhulgakov
Quantization - Dmytro Dzhulgakov
PyTorch
25 PyTorch ONNX Export Support - Lara Haidar, Microsoft
PyTorch ONNX Export Support - Lara Haidar, Microsoft
PyTorch
26 Apex -  Michael Carilli, NVIDIA
Apex - Michael Carilli, NVIDIA
PyTorch
27 Dataloader Design for PyTorch - Tongzhou Wang, MIT
Dataloader Design for PyTorch - Tongzhou Wang, MIT
PyTorch
28 Linear Algebra in PyTorch - Vishwak Srinivasan, CMU
Linear Algebra in PyTorch - Vishwak Srinivasan, CMU
PyTorch
29 PyTorch Mobile - David Reiss
PyTorch Mobile - David Reiss
PyTorch
30 Model Interpretability with Captum - Narine Kokhilkyan
Model Interpretability with Captum - Narine Kokhilkyan
PyTorch
31 Detectron2 - Next Gen Object Detection Library - Yuxin Wu
Detectron2 - Next Gen Object Detection Library - Yuxin Wu
PyTorch
32 Speech Extensions to Fairseq - Dmytro Okhonko
Speech Extensions to Fairseq - Dmytro Okhonko
PyTorch
33 PyTorch on Google Cloud TPUs - Google, Salesforce, Facebook
PyTorch on Google Cloud TPUs - Google, Salesforce, Facebook
PyTorch
34 PyTorch Summer Hackathon Winners - Joe Spisak, Sebastien Arnold, Tristan Deleu
PyTorch Summer Hackathon Winners - Joe Spisak, Sebastien Arnold, Tristan Deleu
PyTorch
35 PyTorch in Robotics - Yisong Yue, Caltech
PyTorch in Robotics - Yisong Yue, Caltech
PyTorch
36 StanfordNLP - Yuhao Zhang, Stanford
StanfordNLP - Yuhao Zhang, Stanford
PyTorch
37 Sotabench for Reproducible Research - Robert Stojnic, Papers with Code
Sotabench for Reproducible Research - Robert Stojnic, Papers with Code
PyTorch
38 Collaborative Natural Language Inference - Sasha Rush, Cornell
Collaborative Natural Language Inference - Sasha Rush, Cornell
PyTorch
39 Privacy Preserving AI - Andrew Trask, OpenMined
Privacy Preserving AI - Andrew Trask, OpenMined
PyTorch
40 CrypTen - Laurens van der Maaten
CrypTen - Laurens van der Maaten
PyTorch
41 PyTorch at Uber - Sidney Zhang, Uber
PyTorch at Uber - Sidney Zhang, Uber
PyTorch
42 PyTorch at Tesla - Andrej Karpathy, Tesla
PyTorch at Tesla - Andrej Karpathy, Tesla
PyTorch
43 PyTorch at Microsoft - Saurabh Tiwary, Microsoft
PyTorch at Microsoft - Saurabh Tiwary, Microsoft
PyTorch
44 PyTorch at Dolby Labs - Vivek Kumar, Dolby Labs
PyTorch at Dolby Labs - Vivek Kumar, Dolby Labs
PyTorch
45 PyTorch Developer Conference 2019 - Panel Discussion
PyTorch Developer Conference 2019 - Panel Discussion
PyTorch
46 Using deep learning and PyTorch to power next gen aircraft at Caltech
Using deep learning and PyTorch to power next gen aircraft at Caltech
PyTorch
47 Named Tensors, Model Quantization, and the Latest PyTorch Features - Part 1
Named Tensors, Model Quantization, and the Latest PyTorch Features - Part 1
PyTorch
48 TorchScript and PyTorch JIT | Deep Dive
TorchScript and PyTorch JIT | Deep Dive
PyTorch
49 Announcing the PyTorch Global Summer Hackathon 2020
Announcing the PyTorch Global Summer Hackathon 2020
PyTorch
50 Opening Up the Black Box: Model Understanding with Captum and PyTorch
Opening Up the Black Box: Model Understanding with Captum and PyTorch
PyTorch
51 PyTorch Mobile Runtime for Android
PyTorch Mobile Runtime for Android
PyTorch
52 Torchvision in 5 minutes
Torchvision in 5 minutes
PyTorch
53 3D Deep Learning with PyTorch3D
3D Deep Learning with PyTorch3D
PyTorch
54 What is Torchtext?
What is Torchtext?
PyTorch
55 TorchAudio: A Quick Intro
TorchAudio: A Quick Intro
PyTorch
56 PyTorch Mobile Runtime for iOS
PyTorch Mobile Runtime for iOS
PyTorch
57 PySlowFast: Deep learning with Video
PySlowFast: Deep learning with Video
PyTorch
58 PyTorch Pruning | How it's Made by Michela Paganini
PyTorch Pruning | How it's Made by Michela Paganini
PyTorch
59 Measuring Fairness in Machine Learning Systems
Measuring Fairness in Machine Learning Systems
PyTorch
60 PyTorch for Hackathons
PyTorch for Hackathons
PyTorch

Related Reads

Up next
Build an AI Voice Assistant with Python | Listen, Think & Speak | Tamil | Karthik's Show
Karthik's Show
Watch →