Continuous Deployment with Weights & Biases Automations
Skills:
ML Pipelines90%Supervised Learning70%Unsupervised Learning70%CV Basics60%Research Methods60%
Key Takeaways
The video discusses Continuous Deployment with Weights & Biases Automations, covering topics such as machine learning model deployment, researcher productivity, and scalability, with tools like Weights & Biases, Chat GPT, and Red AI being utilized.
Full Transcript
thanks for the great introduction uh thanks for having me here uh our friends are weights and bioses are always super helpful for us and also invites us to these things this is my second time talking for them uh so again thanks for the opportunity as said I work for red AI uh this is almost my fifth year in the company uh I'm the first software engineer hire uh I've been through uh many fundraisings you know years of growth uh and I've been part of the model team in the last two years and honestly before that we weren't like crowded enough to have you know that that type of Separation uh so uh I was lucky to work with uh different researchers over the years uh that got that contribute to to many many machine learning models we have that is currently serving Radiologists uh so yeah that's that's a quick introduction about myself uh for from here uh I have three overall objectives today uh two of those which I'm pretty sure I can I can get done the third one is to show you a small pipeline at least the main main modules components that you can use uh to utilize V and bias automations uh maybe I can I can go through that too it you know when I tried this at home it took me a little longer than I wanted but maybe we'll go fast uh for your questions uh let's keep them to let's keep them until the end so I have some time uh some extra time to go through all the material I want uh so that's about it let's uh let's start uh I'll start with introducing red AI uh in more detail uh maybe maybe give you some history uh and what we do and how we are changing how Radiology is being done uh in the United States uh for the last several years uh so uh here's a nice uh visual uh that you know somewhat accurately represents what we do it's B bit more cartoonish but uh that's how I like to imagine uh red AI helping Radiologists uh in real life let's say uh so empowering radiologist with AI and let's get started uh quick background red AI was fed in 2018 uh with investment uh led by Google's AI found gradient Ventures uh they focus on uh these type of SE stage startups and in and deep learning space well back then it was called more deep learning these days it's it's more like Ai and generative I suppose uh primarily uh we have three products so far uh Redi Omni which has been our Flagship breakthrough product uh we have two more new products that are ALS also gaining a lot of traction very fast red ey continuity and red ey reporting uh so today I'm mostly going to focus on red eye omy and what it does and how it change it you know at least slightly change how radiologist Radiology is done uh and it just made many Radiologists lot more efficient and probably make them their lives easier uh by reducing the repetitive work that they end up doing every day when they are at work uh so from there in new Radiology workflow so in here I'll be I'll be talking about our first product Omni uh which you know I've been I've been every part of its its development over years uh so I think to to best explain it I need to give you some little bit of like context or or how Radiology Works uh so let's let's take a step back from the diagram uh you know when you let's say you you you know you you you need to get a radiology study you know think about an x-ray SCT an MRI you know bunch of reasons from you know breaking your arm to having a little bit something a little bit more serious so we all somehow at some point in our lives uh you know go end up in a hospital or had to had to get a scan of some sort uh so the way this works is you would go to you would you would you know you would end up in a location where you can get these studies done uh by a you know technologist you know some some you get your scans get taken uh then that information one way or another this this this varies across hospitals uh or States you know or or or or or general the the the practice but in some way that study needs to end up at a radiologist desk so the desk in M this is you know more like a traditional way of saying it but modern days it's a workstation there's a cube of reports studies that you need to go through uh with some velocity uh through the day to you know to to uh uh to analyze as many studies as they can so that's you the Radiology flow you go in someone takes your scan and that that information those images ends up in a Rous workstation which is usually not the same location that you got your scan at uh you know for many reasons uh and that's usually what happens so radiologist sees the study looks through the images puts in defin things then they would put in the section that is called impressions in the Radiology jargon which essentially from my perspective not really different than a conclusion section for an essay or or an academic paper where where you where you synthesize the previous information in this case would be the findings or or maybe the the indications or some combination of those uh you also list down the implications of the previous information uh which might not be part of the uh part you know part of the existing information but something you can drive from looking at that information and you know some bells and V around that uh so that second part it used to be the case that that second part would be written by a radiologist ER until roughly the year 2020 21 uh so since then uh we built this product uh where the radiologist would put put in the preliminary findings then we would receive that information through our apis we have a machine learning model well more precisely series of let's say inference Services uh that you know take place in generating that final Impressions section of the report so so another simple way of thinking about this this is instead of writing two blocks in a report Radiologists these days usually write only one block and we would generate the template for the second block well radiologist reviews and approves the report the part that we generated uh they they make sure it's it's accurate uh it has high quality and they add any incise if necessary uh but you know we make sure we produce a draft that is that is good enough by itself and and and workable with as is and that's usually the case so I have a small video here uh that would you know demonstrate what we're doing uh better it's it's a little bit hard to follow if you're not familiar with the pieces of software you see in the screen uh so maybe I'll pause and maybe come back and forth but this tries to this this is just trying to show you what I just explained so radiologist you know creating the first section of the report and our our software completing the second part but what's happening is we have a radiology report this is a specialized software uh for radiology writing Radiology reports and our application is actually pretty invisible which is something we really like it's sitting on this uh in this little box here uh and it's it's watching what's going on your editor and from there when it's time you will see that uh The Impressions the conclusion section for the report will be injected magically uh in a kind of a sudden manner following the cursor here okay so it's already in here too uh so that's one example I'm not sure if it's super visible uh maybe let's just look and get let's try look and at it one more time uh this is the shortest example I found but you can see here the impression section is empty uh and the all the information is right about and you see the the user scrolled up and the impression is right there uh again this used to be this information here this impression section this six items would be created by radiologist previously and now we do it uh so it's we have you know we work with at 10 out of the nine largest Radiology groups in the country uh we have more than I guess around 50 customers now so if you received a scan recently it's highly likely that you know you You' gone through this your your study go gone through this process so just to summarize last year we have generated over 15 million Radiology reports uh V around 1. seconds median latency uh and as I said we have many many Partnerships and our products are you know gain a lot more traction uh if you know about radiologist they probably either work with red AI or heard about red AI in the last couple years uh okay so this is the part where I wanted to sort of catch you up with what we're doing uh and I think I'll take a brief pause here even though I said let's keep the questions a little later if you have any questions this is about time if something is not clear let's let's make sure it's clear because from now on we'll go get a little deeper in how this is done and how weights and bias is helping us doing it okay another image that I really like uh you know I got really better at this uh once the chat GPT is out uh so how we do it so I am responsible for de's online inference systems so brief brief summary or definition is when I say online inference the way I think about it is there's usually someone on the other side waiting for that inference right if you think about if you think about the flow I explained here and if you think about that radiologist needs to do this many times many many times throughout the day when they create those findings they would like to receive the impressions as fast as possible so this this this requirement has always been very very Paramount in in in our tradeoffs like you know in our decision making uh and it it really altered the way we think about our system uh it's also slightly different than than the current Uh current you know chat boss you interact with which which take a little bit more time and you are usually fine with that uh you know partly you're not in a hurry in this you're not doing this rep work that you need to go through some volume every day uh partly I guess the illusion of streaming makes things look a little faster and all that uh so from here let's go a little bit deeper in the technicals uh and and the red AI m challenges in making this happen and I'm going to tie to tie to whites and biases from here so the first and most significant part of of our server site is we have we are using machine learning in AI not just for the main task that we are doing so this would be this would be the you know the product I show showed you uh we also use it in many different areas we use it for monitoring our systems we use it for evaluation evaluating our results we use it for customizing our results for the doctor you know our customer support team uses uses services that are backed by Ai and this so on and so forth uh like since from the early days our team was very much skewed towards towards more like scientists and researchers so when we had a problem the usual approach would be can we make a model for it uh which is which end ended up being very p we solved many problems that has never been tackled before we can do many things that you you would have a very hard time implementing with a rule base approached so it's all great uh and wonderful but so the problem here is once you have this type of system uh where different machine learning models are are are sort of central to business processes and like they're that integral right uh what happens is you just need to you end up in this sort of model hell in the sense there would be many many models that different researchers are leading that are serving some critical business process and these models they just need to be updated frequently right they need to be operating with with with with with the quality that you that our customers uh customers expect from us or or a different way to do is just you want to put your res Searchers in a position to build more models you know change them more frequently ship more models uh so this has been a challenge that we've been dealing with since the early days uh these machine learning models come with very specific resource requirements to run them uh you know they all have they all need this this data to be pulled in whatever is going to run that model so they all have they have these sort of peculiar requirements uh uh you know puts you in a position to come up with different solutions let's say so this slide it's more like I'm going to talk about our approach General approach so when we sit down and we we think we thought about okay what are the properties we desire from a server s side system uh so that we can do this we can build this product and you know keep iterating as frequently as we want and you know create ourselves this this this this this this this space in the field uh the first the first thing that is obvious is and this is probably shared with many any company it's frequent deployment and it's it's getting quick feedback so this by itself uh requires a streamlined approach as as you want to update the existing model collector training data as fast as possible even though it doesn't make sense at all times you want to be in a position to do that when it matters the other approach the other iCal Point Property we desired from from our system is obviously the researcher productivity uh usually when it's when a model is going to be part of a production system and it's going to be serving customers there will be a synchronization step with researchers and and machine learning operations professionals let's say or like or the engineers on the team that are more more closer to software side uh and we find that that synchronization is also has a potential to slow researchers down uh so we wanted to get around that too and the last part is as I mentioned our systems are online which means that they usually needs to be up 724 and they need to maintain an operational quality uh throughout that time period that they're up uh this also this requirement if you think about it we're also integrated in the healthcare workflow you know radiologist workflow and it there is also I think uh sort of another another element of it that that makes this operational quality uh important as as you're part of the healthcare flow and once luckily there only has been a one brief aage in the recent years even then when that happens Radiology slow down so just overall our capacity to read Radiology reports in the country goes down suddenly and you know very undesirable uh problem aside from it being a simple B problem I would say so again frequent deployment researcher productivity uh scalability operational quality these are properties that we can't really give up if we give these up our product wouldn't work so these would be these were the these were the properties that we we were seeking for uh now going back to weights and biases so weights and biases I think if you have been using the product for a while it used to be more like a research oriented uh you know training oriented product where you would you follow your experiments you know you you can see your you know loss functions converging that's a you know great thing to see and it's it's it's how you how how we how we adopted V and biases I kind of remember the day someone reached out to me hey can you add API keys to this this this instance something like that that's how it started I think this was 2000 late 2021 20 2020 early 21 that's that's how I remember it again this I think this part the model register you're all lot more familiar than what I'm going to talk about this is where when you train a model you just lck some intermediate results you know we usually metadata uh you know how long things you think what is your loss all things like like so we have been using model registry you know at least three four years and it's been around for for for a long time and it was our you know it was I think it still is every you know it's everyone's favorite tool especially for researchers and deep deep learning you know experts uh but what I'm mostly interested in is the weights and bias automations uh so you know two slides ago I was talking about the properties that we're looking for uh in a in a system and to make some of those happen uh last year around these times I was in a search uh of of like a tool or or like a sort of a software architecture uh to address some of the some of the goals to achieve those properties uh that I mentioned so let's maybe go over some of them really quick uh so one one one one desire I had was to implement an infrastructure where these models are changed in an event driven manner so as I mentioned there was a lot of synchronization required between researchers and software developers uh like more like software engineers in the team to deploy these models uh and that involves a lot of repetition and a lot of human labor you know this repetition can come from just writing the same code or most of the same code over and over because you know it's it's not centralized and it's it's not being reused properly or it's just human labor you know people talking to each other hey I I have this model oh you really have that model when it's going on it's going on next week oh I didn't know well it was on the schedule you know some some these type of discussions that you want to eliminate as soon as possible uh the other was automating the change proposals where when it comes to changing infrastructure with the new model is it's it's not really obvious who's supposed to do it or who's supposed to own it ideally these new models are made by researchers so if they can propose the required changes to your infrastructure somehow even though they don't know how exactly it would work or what needs to be proposed as long as they can trigger it we figured out that we can make it happen so the the the third one is again a very simple property I'm pretty you know prettyy much every server side engineer would would like to see a normalization like a process normalization uh but we needed this to just you know make the product work we needed some normalized consistent processes across models and those processes allowed us to sort of reduce the process variation across models uh and you know put us in a position to sport all all those research bandwidth and deploy all these models in a frequent Manner and deliver improvements by responding to customer feedback is they interact with our models so this is how we use automations this is uh now I'm going to go deeper into the actual steps where we use automations and how they how they help us I'll also talk about a broader uh sort of a broader continuous deployment workflow Maybe uh so this this image just you know lacks completely lacks rigor and it kind of looks silly uh but here my defense for it so whenever I think about a continuous deployment pipeline uh easily when you implement these they get increasingly complex there's many BS on missiles different components talking to each other uh so I usually start with this type of approach where I I think okay what needs to happen in a machine learning deployment workflow so this is how I think we need to train the model we need to upload the model somewhere that and then you know that model needs to get registered it needs to be somehow you know stage for deployment then it needs to get approved by Machine learning operations Engineers or software Engineers on the team that are responsible for this then it needs to be run so these this is how I think about it and it's this is very simple as I said and it kind of looks funny uh but the next slide will be more satisfactory to to you know technically savey uh so this is the this is a very much stripped down version of an pipeline we use internally actually uh I took out you know some some bells and whistles that wouldn't matter overall in the in the in the outcomes uh as you can see on the left hand side we have the standard process of model training I'm not going to go into that in this talk uh you know it's it very much depends on your your own company your own use case how you go about it you know you can use uh you can use a distributed training cluster on the you on a cloud you can just use a single instance get somewhere that's not part of my talk but in the end one way or another once you have a model uh your model will be in the form of some data some artifacts you know these can be weights some configuration files things like that I'm sure the hugging face models that you know you see uh are a good example of this there's a there's a normalized model export uh so this step where where this step gets relevant in deployment is the bottom part once the models are the the artifacts are uploaded they are linked to weights of biases model registry so there would be an artifact record created in Wast and biases and that artifact that art AR fact record will have a pointer for these this S3 location so it can be picked up from there this is purely built around weights and bu automations and I think the most critical step here is the first step once the model is registered to rates and biases the artifacts are on S3 the next step is if you read through the automations documentation is adding an Alia to an artifact when you add that alas on an artifact that you just registered weights and biases triggers a web hook that makes a post request to trigger a get up actions workflow and then that get up actions workflow automatically looks at what is in the current infrastructure what will be the new state of the infrastructure after the change proposed let's say if you were to if you were to deploy this newer model replacing an older model then it will open a pull request for you so the people responsible for the platform people the software Eng responsible for running off that model can review and approve so two I think two very interesting points about this that you know uh the first one was it took me a while to get used to that idea is just automated P request normally you know you're more used to Developers opening pool request and writing code so this utilizes templates uh to to create terraform modules and open pull request against the main branch uh where where the infrastructure is synced with uh the we use we use terraform Cloud to orchestrate these INF infrastructure changes and apply these infrastructure changes usually this works in a way that when there's a pull request open in your repository that you declare your infrastructure terraform cloud or or Atlantis or or many many many continuous deployment integration software you know watches these events in your in your repository uh then can then can create these plans or apply the infrastructure changes uh so then this work this pipeline ends up getting an approval for that automated P request by an you know mlops member uh that has hopefully some familiarity with your infrastructure and can can approve those changes to your infrastructure uh which in the end these days we use Sage maker inference endpoints uh but we would end up either replacing an existing endpoint or or spin up a new Sage maker inference standpoint uh so again as I say two two major points here the first one is the fact that the the P request is written by a machine and not usually something you do and the second one is there is no there is no programming involved there is no developer communication involved all you need to do is add an alias to an artifact and then the the the staff responsible for the infrastructure reviews the PO request that this that this Alias triggers well the this this get up actions op let's say so let's see I think I'm right about right at time right now yeah uh it's took quite a little late though but anyways uh so the next would be to go over the implementation uh I'm not sure can I do like five minutes three minutes okay two three minutes so uh I'll I'll make sure to share this repository I actually have it with me here uh that has the required tooling for you to make a simple example of what I explained here uh I have a well the repository has a script that you can log an artifact and Link it to a model there is a GI up action workflow that handles these weights and bias automations events that comes after when you add an alias there is a CLI for GitHub actions to use so this is this would be a Python program that uses Ginger templates to generate these terraform modules and you know some some some some operations to uh create that P request you know committing the change pushing that change opening a here all these are programmatic automated and the last one would be the actual terraform module where you announce announce your infrastructure so this would be you know the actual actual concrete infrastructure components say in your Cloud environment let's say a this can be an AWS AG maker point or an S3 bucket or whatever else and I am out of time if I oh okay I want to I want to highlight this this slide this is the most important slide anyway so uh we are actively hiring if you are interested in what we're doing if you're interested in improving Healthcare in the United States if you're interested in helping Radiologists if you're interested in just machine learning and working with us uh find us scan the link find us after the talk uh we'll be here you'll be surprised there's actually four four more people from Red AI sitting in the back any wood work all of them are friendly most of the times you know uh please do that and if you're interested in seeing the implementation and the pieces uh I'll work with the weight and biases team to you know make this public right now it's hooked up to our companies the system so you know uh maybe not the best idea to make it public right away uh I can work on that with them because I'm out of time but at least I you know I was able to deliver the the slides uh that I wanted to uh hopefully I introduce you guys to Red Ai and you have a better idea what we are doing hopefully you got a little bit interested and thinking about applying and talking to us
Original Description
👀 *All of the Fully Connected San Francisco 2024 videos are available at http://wandb.me/fcsf24yt*
*About Ali Demirci's Session On Continuous Deployment with Weights & Biases Automations*
In this session from the Fully Connected conference in San Francisco, Ali Demirci, Senior Machine Learning Engineer at Rad AI, delves into "Continuous Deployment with Weights & Biases Automations."
Ali discusses Rad AI's journey in enhancing radiology practices using advanced machine learning models and automation tools. He covers the critical aspects of frequent deployment, researcher productivity, and maintaining high operational quality. Learn how Rad AI is setting new standards in the medical field by improving radiology report generation and overall workflow efficiency.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Weights & Biases · Weights & Biases · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
0. What is machine learning?
Weights & Biases
1. Build Your First Machine Learning Model
Weights & Biases
Intro to ML: Course Overview
Weights & Biases
2. Multi-Layer Perceptrons
Weights & Biases
3. Convolutional Neural Networks
Weights & Biases
Weights & Biases at OpenAI
Weights & Biases
Why Experiment Tracking is Crucial to OpenAI
Weights & Biases
4. Autoencoders
Weights & Biases
5. Sentiment Analysis
Weights & Biases
6. Recurrent Neural Networks [RNNs]
Weights & Biases
7. Text Generation using LSTMs and GRUs
Weights & Biases
8. Text Classification Using Convolutional Neural Networks
Weights & Biases
9. Hybrid LSTMs [Long Short-Term Memory]
Weights & Biases
Toyota Research Institute on Experiment Tracking with Weights & Biases
Weights & Biases
Weights and Biases - Developer Tools for Deep Learning
Weights & Biases
Introducing Weights & Biases
Weights & Biases
10. Seq2Seq Models
Weights & Biases
11. Transfer Learning for Domain-Specific Image Classification with Small Datasets
Weights & Biases
12. One-shot learning for teaching neural networks to classify objects never seen before
Weights & Biases
13. Speech Recognition with Convolutional Neural Networks in Keras/TensorFlow
Weights & Biases
14. Data Augmentation | Keras
Weights & Biases
15. Batch Size and Learning Rate in CNNs
Weights & Biases
Applied Deep Learning Fellowship Overview and Project Selection with Josh Tobin (2019)
Weights & Biases
Grading Rubric for AI Applications with Sergey Karayev (2019)
Weights & Biases
16. Video Frame Prediction using CNNs and LSTMs (2019)
Weights & Biases
Image to LaTeX - Applied Deep Learning Fellowship (2019)
Weights & Biases
17. Build and Deploy an Emotion Classifier (2019)
Weights & Biases
Applied Deep Learning - Data Management with Josh Tobin (2019)
Weights & Biases
Snorkel: Programming Training Data with Paroma Varma of Stanford University (2019)
Weights & Biases
Applied Deep Learning - Troubleshooting and Debugging with Josh Tobin (2019)
Weights & Biases
Troubleshooting and Iterating ML Models with Lee Redden (2019)
Weights & Biases
Designing a Machine Learning Project with Neal Khosla (2019)
Weights & Biases
Lukas Beiwald on ML Tools and Experiment Management (2019)
Weights & Biases
Building Machine Learning Teams with Josh Tobin (2019)
Weights & Biases
Pieter Abeel on Potential Deep Learning Research Directions (2019)
Weights & Biases
Testing and Deployment of Deep Learning Models with Josh Tobin (2019)
Weights & Biases
Five Lessons for Team-Oriented Research with Peter Welder (2019)
Weights & Biases
Applied Deep Learning - Rosanne Liu on AI Research (2019)
Weights & Biases
Making the Mid-career Leap from Urban Design to Deep Learning/Data Science
Weights & Biases
Organizing ML projects — W&B walkthrough (2020)
Weights & Biases
Brandon Rohrer — Machine Learning in Production for Robots
Weights & Biases
Nicolas Koumchatzky — Machine Learning in Production for Self-Driving Cars
Weights & Biases
My experiments with Reinforcement Learning with Jariullah Safi
Weights & Biases
Applications of Machine Learning to COVID-19 Research with Isaac Godfried
Weights & Biases
Testing Machine Learning Models with Eric Schles
Weights & Biases
How Linear Algebra is not like Algebra with Charles Frye
Weights & Biases
Predicting Protein Structures using Deep Learning with Jonathan King
Weights & Biases
Rachael Tatman — Conversational AI and Linguistics
Weights & Biases
Reformer by Han Lee
Weights & Biases
Sequence Models with Pujaa Rajan
Weights & Biases
GitHub Actions & Machine Learning Workflows with Hamel Husain
Weights & Biases
Look Mom, No Indices! Vector Calculus with the Fréchet Derivative by Charles Frye
Weights & Biases
Jack Clark — Building Trustworthy AI Systems
Weights & Biases
Surprising Utility of Surprise: Why ML Uses Negative Log Probabilities - Charles Frye
Weights & Biases
Track your machine learning experiments locally, with W&B Local - Chris Van Pelt
Weights & Biases
Antipatterns in open source research code with Jariullah Safi
Weights & Biases
Attention for time series forecasting & COVID predictions - Isaac Godfried
Weights & Biases
Made with ML - Goku Mohandas
Weights & Biases
Angela & Danielle — Designing ML Models for Millions of Consumer Robots
Weights & Biases
Deep Learning Salon by Weights & Biases
Weights & Biases
More on: ML Pipelines
View skill →
🎓
Tutor Explanation
DeepCamp AI