From Infrastructure to Application: Lessons in Building Scalable ML Systems
Key Takeaways
This video discusses lessons in building scalable ML systems, covering topics such as infrastructure, data science, and the evolving landscape of machine learning and AI systems, with a focus on LLMs and their limitations.
Full Transcript
of 2025 I am so excited to be here today with faras hammad hey faras what's up man hey what's going on here you go great to be here I'm I'm so excited to jump into talking about infrastructure applications how we actually how infrastructure and data SCI and ml can actually deliver um business value particularly in Building Systems and not just models um really excited to talk with you especially with your you know breadth of experience being now being a machine learning leader at door Dash formerly at Netflix meta Uber and I didn't know this but before that your first job um where you started doing AIML stuff at Yahoo as well working on personalization there yeah yeah I did the full Bay Area tour um you know got got like my uh my punch card uh just maybe like a couple more more left in me fantastic um I'd like to welcome everyone who's joining live um welcome to out of bounds fireside chat the our first one for 2025 uh please introduce yourself in the chat let us know where you're watching from what your interest is do you work in data science ml AI um and and and what are you up to um a bit about outabounds out of bounds uh builds open source software and products that help you build both ML and AI applications including llms um out of bounds works on infrastructure and productivity tools that allow data scientists and MLS to focus on building models and systems and doing science while having easy access to infrastructural layers such as compute orchestration versioning um doing a lot of that through the open source metaflow but please check out out of bounds platform as well and I'll link to the GitHub for metaflow uh in the chat um and it isn't always that I get to do this um but I'd love to hear your thoughts on metaflow for us and maybe you can tell us what metaflow is useful for having worked on it and worked with vill of course at at Netflix as well yeah I mean you know full disclosure uh I did work on uh uh on metaflow at Netflix uh and was part of the team that that helped open source it uh so my view may be a little bit biased but you know it is one of the the best software products that that I have used um it's I guess from the uh um you know outside looking in it's deceptively simple right that the amount of things that it handles really makes uh the ml ml's life a lot easier uh and it's it really hits that sweet spot in terms of abstracting at the right layer right so it's opinionated where it needs to be but um where uh probably mes have uh care the most about it really lets them uh it really gives them free reign right so that's yeah that's that's really the the distinguishing factor that I really see with met you know it it's built in such a way that it's easily integratable into a variety of different Stacks right um and even if there isn't an open source support for you know whatever you're depending on internally it's pretty straightforward to build your own custom plugin um but more likely than not you actually find somebody has already built that plugin for you so um you know I'm I'm a big fan of it yeah awesome um and in terms of the abstraction L provides I I do agree one example I like to give is the at kubernetes decorator and this is one of one of many but you know as I'm I'm I'm a scientist man I'm not a software engineer or infrastructure engineer so the fact that once I have a cluster provision for me I can just slap on that decorator specify the memory what whatever I need if gpus or whatever and attach that to a step and it just it does the work send it up there and bring the results back right yeah yeah it works and it's reproducible right so if yeah you yeah if you if you run something once uh it will run in a different environment and others could run it as well right so that a huge thing um a lot of software Engineers take reproducibility for granted but that's one area um you know from from the ml side that is lacking and and you know at a lot of companies and and and you know hasn't become as standard practice as it should be totally totally well let's jump in man and to your point I I think a lot of the standard practices that have developed in like more Tech forward like serious tech companies like the ones you've worked at part of the work now is to get all of those behav avors and and systems and let people know what what's happening right which is one reason I'm super excited about chatting publicly with you about this because you've worked at Netflix as I said meta Uber door Dash um and Yahoo so I'm wondering like what drives your passion for machine learning and how have these experiences in in their different way shape what you do in your approach yeah yeah I mean the what got me interested in ml is the same thing that that got me interested in in computer science like more broadly uh ml is a space that's applicable to almost anything um really provides the opportunity to pursue both uh depth and breadth right if you want to go deep in a specific area you could go like for decades right if if you want right but then there's also just so much bread to it as well right uh different areas of uh of ml uh AI um you know everything from from knowledge graphs to LMS to to whatever um and then you know the ambiguity that that comes along with with really productionizing ml systems really gives you the opportunity to to express your creativity right so like how do you model the problem what proxies uh can you use to get like ground truth data what are you're optimizing for you know just all these different aspects um you know is is is really what drives my passion for it amazing um and suppose um before moving into kind of how ml can deliver value to all of these different types of companies I'm wondering at a place like door Dash Why is ml infrastructure so critical and what impact does the infrastructure have on business outcomes yeah uh so if you look at door Dash um really the the breadth and variety of problems that uh require ml or want to leverage ml right you have Store recommendations we have tools that used by uh Merchants improving search uh you know our product knowledge graph there's some interesting blog posts related to that if you guys want to check out the door Das plug just a quick plug there um even like uh core info right has used cases that that want to leverage ml right so you have the diversity of problems and we're also very early right in our in our um you know life cycle uh and we're looking for step fun uh step function uh uh growth or or leaps even in the business right and that requires uh a lot of innovation and infra is really what drives Innovation right um if you want to uh you know really make uh an impact in in these you know almost Green Field areas you need to have infra that's set up to to well well support that right to be able to quickly uh jump in in a new area try out different things and that's one of the reasons why ml INF specifically at door Dash is is so critical to the business and and ensuring successful business outcomes fantastic and I've linked to um the engineering blog for do Dash um so people can definitely check it out I am interested um you know we live in a world where llms reign supreme in the cultural Consciousness at the moment and you know we're told this is the year of Agents um you've mentioned knowledge grafts a couple of times and I think this is this is something that we all should be talking a lot more about um and before we started the live stream of course you told me about some knowledge graft work you were doing at Yahoo in the early days as well so maybe you could just briefly tell us a bit about knowledge graphs and why they're so useful yeah I mean I built knowledge graphs both at Yahoo and Uber actually um and in terms of uh when you actually talk about about AI right like the broadest uh concept of AI is actually a knowledge base right um that's where artificial intelligence technically starts right from like a you know academic definition uh and with uh you know the concept of knowledge graphs you could really uh bring in data from a bunch of different data sources uh make connections that you normally wouldn't make uh and even derive information based off of these existing uh definitions that don't exist anywhere else right um so like the canonical example we used to use back um in like the Yahoo njra days is uh you know you know that um Malia is Barack Obama's daughter and that Sasha is Malia's sister and from that you could you know uh essentially derive with you a certain level of confidence or probability that Sasha is also um the daughter of of Barack Obama right uh you didn't necessarily read that information from any any Source or get that information from any Source but you essentially were able to derive a fact right um and basically making those uh uh connections uh is really helpful you know um when it comes to uh you know in that case it was is force supporting search but you know know a variety of different Ai and ml applications I love that you gave that example of family relations in in the Obama family because it actually reminded me of something that um llms can be completely horrible at right um and there is the example I I can't quite remember it but it's if you say who is Tom Cruz's mother to chat gbt it won't give you the correct response but if you say who is and put her name there it will say that's Tom Cruz's mother so the tokens are linked in a very in in in a very strange way right so we don't have the the robustness of systems such as knowledge knowledge crafts right yeah yeah and you know like LMS when it comes to um giving factual information I think like the there's well documented issues with that right and if you look at how they're trained and built they're actually not really even trained for that right it's it's more of like a you know probabilistic type uh type deal in terms of predicting the next token yeah not to your point so they're try I mean if we're building a probalistic next token prediction system why would we expect it to be accurate on top of that I think with rhf4 DPO we've got systems which are reinforced and optimized to seem helpful as opposed to be accurate right so exactly exactly there designed to sound correct not be correct exactly and you know that's something and and that's how they are now you know maybe in the future uh uh the architecture evolves in in a in a different way yeah uh you know more likely I I think we'll probably see um almost like a Ensemble uh type approach where you have LMS alongside other model types uh that try to to correct um or catch errors um of whatever that LM is outputting right and you know depending on the cost of compute in the future you know maybe that's more viable right very much you could actually I've done this so you know I do a lot of online interviews and chats like this and then I want to get verbatim um quotations from the transcript I want chat gbt or clae or whatever to help me do that it hallucinates all the time so I've built a system which will get it to a bunch of interesting things then I use fuzzy matching and string matching to actually get those out of my documents right so yeah exactly and like that that's something that we just need to design for and factor in right if you're trying to leverage LMS um you basically have to factor in like it's going to be wrong a lot of the time right like don't treat it like an oracle um just yeah and don't think of it we need to change our mental model around what software is as well it's non-deterministic on top of that I love that um one of the first hello worlds for a quote unquote agents is get an llm to call a calculator and execute python code or something like that um and pointing out that you know these things aren't great calculators which is something we would expect of software right yeah yeah yeah for sure for sure um so let's step back a bit and I'd love to hear about your your experience es and I'd love to kind of frame it in thinking through what what are the biggest differences in how companies like Netflix meta Uber door Dash Yahoo approach ML and AI infrastructure and applications yeah yeah and um I guess like looking at the the most recent you know few companies Netflix meta and Uber um uh they're very different culturally uh you know at least by Bay Area standards right so Uber at least at the time uh was very operations driven right so we had a lot Ops folks they ran the city teams they they did amazing work uh and it was all about execution getting something out the door then rapidly iterating right Netflix uh culture was um uh was you know is is pretty unique uh in terms of keeping teams small uh and specialized hiring right so they hire people for a specific role right there wasn't any uh generalized hiring which a lot they areas Bay Area companies uh at least used to do um that you know you were hired for a specific role or task and you know you came in with you know very strong opinions about what tools you want to use and how you did things uh with meta it's just kind of like the sheer scale is almost what makes it unique and not just in terms of like the number of Facebook and and Instagram users but the number of people working in the particular domain right so you know Facebook has what over 70,000 employees like 30,000 engineers and you have dozens if not hundreds of people proposing changes to one model right so all of that is of course manifested into how they build their infra and and ml infra right so and the and you know of course like Yahoo um that was so long ago that it was like before a tensor flow and kind of all these things came into play that speaking about that experience it it'll almost be like describing uh like floppy discs or something to to sound the way I maybe even stone tablets no I'm half joking there but um the story you told me before we started was so wonderful would you mind sharing sharing that yeah yeah I mean I just remember like uh you know I got my first taste of uh uh ml uh working on you know homepage uh personalization and I do remember like you know essentially implementing linear aggression offline and then like manually committing the weights into some PHP code and like Computing that online to dynamically rering components of the page like you know that's what we did right uh you know not all of yaho did that um that you know uh for for certain use cases that that was the level we're operating at um it was also all about VMS um you know when we're talking about like the the classifiers that went into the knowledge graph like determining you know what class this entity supposed to belong to or this person belongs to or you know what two items are duplicates um I can't remember the last time somebody mentioned svms to me right without a doubt yeah yeah but I guess going back to uh the other three right like looking at at you know Uber and and Michelangelo specifically right um so Michelangelo is uh Uber's ml platform uh and if you look at the you know the first few years of it of its existance um they took a very constrained approach right so all model training was actually done via UI uh and you basically were able to select from three different model architecture types so you could do a tree based model logistic regression and I think it was like lstms and I'm not even sure if lstms even worked M um and you know you select you know a few predefined features Etc and you could maybe change six hyperparameters and then you hit train right um so very constrained but you could get things out very fast right and you didn't necessarily require um too much domain exper uh expertise in in ml uh it was more optimized for people with domain knowledge in their their own area right so if you KN knew pricing well if you knew your your your city well right you could select the features that you think are are most relevant for your particular task and you know what don't worry too much about the the ml side you know we'll put up these scar grows for you Netflix took a completely opposite approach right like if you tried to pitch that um at Netflix you'd have a Revolt of all the the um data scientists and and MLS right because there are too many guard rails for them or yeah like you could only you know select from three model types right um and you know forget you know picking your own you know uh type of modeling uh uh Library framework right like tensorflow or pytorch whatever uh you know internally uh metaflow even supported R right for I remember yeah um uh so what Netflix Netflix focused on was allowing um data scientists to do what they would they do best right so being very open to like all right you should be able to use whatever Tool uh you have at your uh that you know you're most comfortable with where you could do your best work that you know what we'll abstract away the infra right um and we have infra Engineers who can do infro right so that was the the the driving philosophy there right um everybody you know do what you're hired to do uh and then meta um again it's around how can you get so many people to contribute to the same goal um without it being total chaos right so it it's almost like a a factory type approach to be honest I didn't really like the the the developer experience right of like trying to um you know modify uh um or like propose new new features to an existing model even so things that might you know would have taken me like less than half a day at Uber would take two weeks uh at meta right because you have to go through so many checks and balances and things like that just to get something done but you know there is no way it could operate in the same way that that Uber operated right just given the number of people working in the same domain that makesense sense um so I'd love to now jump in and think about the role of generative AI in everything that's happening at door Dash and other places today so llms are kind of they're becoming widespread at least in the cultural Consciousness some in production as as well right but they're clearly not a solution for every problem so I'm wondering in your work how have you seen them fitting into traditional LL into traditional machine learning systems and how should companies approach integrating llms into these ex these systems yeah I mean great question so uh I think like we alluded to a little little bit earlier uh you have to take into account that you know llms aren't like an oracle right they are wrong uh significant amount of time right so they have a very high uh error rate not just in terms of like maybe giving the the wrong responses but maybe even the format is is not even what what you might expect uh there's a lot of uh ambiguity there right so um I guess like that's that's the first step really is you know if you're trying to integrate llms uh into a product or or Surface uh you need to understand the the limitations uh for where we currently are with l right um so in some areas uh it's okay right um if the the cost of an error is negligible um so you could imagine you know maybe you're using llms to quickly add uh tags to to restaurants uh or something like that right where it's very user visible the cost of maybe mislabeling one item you know it's not a huge deal right depending on like how bad you mislabel it or how important of it is that uh you know if it's like the longtail items that you're missing labels for that's okay right uh and it's a very rapid way of getting up and running right as opposed to creating your own custom model that you had to train to you know predict tags on um you know these longtail items uh and even that like is going to have an error rate right assuming you don't have like the ground truth that's necessary uh so really just kind of being very clear about those expectations um and I guess like uh overall um I do see uh llms and more kind of traditional models coexisting right uh you could also you know view maybe some more traditional uh models acting as guard rails to the outputs of of of some llms um or using um the output of llms as kind of uh input features um to kind of like more more robust models or or or more traditional models without a doubt and I do I love your framing of thinking through the types of mistakes they can make and how they they will like probabilistically they will be making a certain there will be a certain error rate right and I do think I actually like Andre Kathy's take where he said something he tweeted something like if we're going to use the term hallucinations we should probably acknowledge that llms actually hallucinate 100% of the time and sometimes those hallucinations coincide with ground truth yeah I think that's that's actually right yeah um and he Compares it to search which hallucinate 0% of the time with respect to you know the the ground truth in in the Corpus um I also like love your example of low risk um where an error may not be that costly um having said that in that case you still want guard rails to prevent you know labeling a restaurant with a cuss word for example right yeah yeah anything user visible right really updates um uh UPS the risk uh and you know even if it's not very user visible if it's like on a very popular item right uh you don't want to exclusively rely on llms yeah it could be a great fallback right if you don't have anything else call an LM see what it thinks right or use it as an input um so like the LM thinks this uh let's feed it into a model with some other features and maybe the model could give you like a a better answer yeah right absolutely um as part of this conversation I'm I'm I'm really interested in infastructure um and as you know like metaflow considers you know has an opinionator but I take that really makes sense to me on what the infrastructure stack looks like where you you know you have data at the bottom um you need to vergin things you have an orchestration layer you have um your modeling layer all all all of these things as part of the infrastructure stack right um and I'm wondering in your mind well deploying generative AI at scale often requires rethinking infrastructure in some ways so I'm wondering what you see as the biggest changes to the stack and or significant challenges you've seen in actually productionizing llm powered applications yeah yeah and like this is uh sometimes um you know it's it's very dependent on you know how comp like existing companies implemented their stack right so some companies made the decision or I guess some uh platforms made the decision that like uh they only serve models that have been trained on their platform right and run off you're at a massive disadvantage because you're not built to pull in kind of like a bring your own model type uh a solution right so one thing that I'm actually very excited about in in the last few years that's changed um if you if you go back like you know five five years right even not even that long ago uh if you wanted to do anything in ml you had to be at a large company right they had just the massive advantage in terms of not only their compute resources but essentially the data right but what's changed in the last few years is that open source is actually outpacing uh the traditional big tech companies right um and I think that's the trend that's here to stay uh given how you know really kind of the incentive structure within all these large companies it's very hard to uh and just really just looking at the disparity in terms of the number of people right it doesn't matter if like you know how well capitalized these companies are they have a a finite number of uh people working there um with finite bandwidth whereas if you look at the open source Community it's like orders of magnitude more people and they're just constantly trying out different things right so it's going to be very hard for uh these companies to um outpace uh open source if they just exclusively r on um their internal tools and and products and and malls that they've trained right they really need to be able to Leverage The the open source Community uh so if your stack is not able to easily pull in you know the latest foundational models or if that requires a lot of effort uh that's going to be like a nightmare for you to solve right um if if uh your your stack is is implemented in in a particular way uh the other challenge uh really comes into um really just like the call pattern has completely changed right traditional ml there's usually like a very standard call pattern where you know the canonical example is just like you know recommendations or ranking right so you have like a piece of content or item or store whatever or a list of them and you throw that at um a model that you have deployed somewhere and then you know uh the serving layer calls the feature store and it it passes all that information along and and then you get back a store right so it's a very transactional type type deal but when you know we talk about the llm world like one there's really no concept of a feature store anymore right there's like you know rag based architectures and context windows and things like that but that's not really the same thing right or or the same pattern um uh and people also want to use LMS in a more creative way than you know traditional models right so if you have a if if your stack isn't built to really support that that creativity or customization that's also going to be a massive challenge for you um and I would say like that is probably going to be uh an area that troubles a lot of uh ml platforms across the industry because most of them are built around that that hyper optimization like how can we serve you know millions of of QPS uh you know you basically had to add in some constraints right and constraints need to be blown away totally I've I've never actually thought about this before and maybe a really silly idea but I I totally agree that we don't have feature stores anymore but I wonder if there's some future of like promp management stores and data set stores for fine tuning with with models and what what type of shared resources evaluation stores and and the the the these types of things what and what type of shared resources across an organization would make sense because feature stores arose because everyone was building features on different teams and guess what some of the features people built other people needed voila right yeah yeah and I mean it's it's already a read a need right now right so if you have kind of mult uh multiple people interacting with a specific model they all have different you know context windows that that they care about different prompts that they had um you want to be able to load that in and out quickly um you know those are all kind of changes that that need to be made um at Theo platform level um and you know it is fairly different from you know at least for feature stores it's somewhat similar to a lot of other parts of you know traditional software engineering right if you know you had like certain key value stores things like that you know maybe uh feature stores are hyper optimized and in uh for specific uh query patterns and you try to group features together in in specific ways but it's not too different in terms of a query pattern right you want to load some data and you know pass that in whereas you know loading different context windows and prompts and things like that is a little bit more custom right yeah um totally so how do you see the role of traditional ml models like well no longer um support Vector machines thank goodness but like tree based methods or two tower architectures for rexis evolving alongside llms and if you're able to speak to anything that happening in product at door Dash I'd love to know but if not that's totally cool as well yeah I mean like just kind of generically speaking um you know even even before LMS right uh you know you you a lot of the industry was on on tree based models and then you know dnns became more popular more powerful uh that people still use tree based models right it really depends on on the use case that you have right just because you can use an llm for something doesn't mean you should right LMS are of magnitude more expensive um uh to run and deploy compared to uh tree based models um and they also seem easier to work with on the surface of things but the complexity they actually introduce is incredibly costly not just in terms of the cost of pinging apis or self-hosting models but in terms of the workflows you need to develop yeah workflows and also just like in some use cases you need a high degree of explainability or interpretability right that goes away with the llms um uh and even in in some very specialized use cases right if it makes sense for you to like uh like business sense where even like a small changes in um you know Au or or performance metrics uh matter maybe it makes sense to have like ml go ahead and build out like custom models using traditional methods um than leveraging an LM LM might help you bootstrap in a particular area but you know like going back to the example I used before around um you know labeling uh products um you know you could start with an LM uh maybe you find tune the existing llm but maybe you actually get a great source of of ground Truth where you could train a more traditional model and it performs better than an llm right um so um and yeah uh going back to again dnns versus treebase models if you go to any company um you'd find that a majority of the models are actually tree based b or or very simple models right um uh and for a lot of use cases that that's like good enough right um and they care more about latency than um maybe like even correctness uh so I I definitely see uh continued use of of both um LMS and and more traditional techniques awesome um to your point as well I so I teach a bunch as well and teach data science machine learning AI stuff and a lot of people come in and they're like I want to learn pie torch and ra ra I'm like of course that that's cool but I I do tell people and I think this is something I stole from Jeremy Howard in his fast AI course that if you can do use tree based methods build some deep learning architectures and do some linear and logistic regression that's going to get you 95% of the way for most most things right yeah yeah absolutely so last time we spoke you mentioned machine learning tools um are converging in some ways to serve both Advanced users and non-sp Specialists and I I love this I'm wondering what this looks like what's driving this trend and what what trade-offs are there yeah uh I mean in terms of like what's driving it uh there's a lot of factors like even before llms that um in the current cultural moment that we're in there's just so much interest around machine learning that everybody wants to try to use it right even people without the the same traditional background um and uh as it becomes easier to run uh traditional um uh I guess ml models uh even uh different parts of uh you know um organizations companies Etc are parts of the stack want to leverage ml right so even if they don't have deicated mle support or data science support so how do you get started how do you quickly bootstrap right even like terms of uh uh you know core infra use cases right how do you um can you like build a model that you know predicts if a node is about to go down right or how long a particular job is going to take to run right for most companies they're probably not going to hire an m just to solve that one use case right it might you know might have significant um impact in terms of uh infrastructure cost um and kind of overall reliability of the system right so the software Engineers on that team would probably want to be able to quickly bootstrap uh such a solution uh validate whether it's viable and maybe later on they get more investment from you know the the centralized Mily org or they're able to hire a data scientist to to actually build a more robust model if you're just trying to establish a baseline um you want like a a quick and dirty way to uh to really evaluate the the use of ml in your system um so we've been talking around something and it's the relationship between infrastructure infrastructure Engineers data scientists machine learning Engineers platform teams um and I just want to like talk explicitly about it so essentially what we've been talking around is that ml work in like depends heavily on collaboration between a variety of different teams with a variety of different skills um finnally one thing I I really liked when I last year I went I was in London with Savin Goyle the CTO of out of bounds um who worked on meta at Netflix amazing guy and we went to Bloomberg's office together in London which is one of the wildest places ever um it's it's incredible but one thing I really loved is that in the office the infrastructure team sits next to the data science team so there's just like incidental conversations and Vibe flowing back and forth that just helps collaboration um so much I'm wondering what strategies you've seen for fostering better collaboration in these types of cross functional teams particularly infrastructure than people who use the tools yeah yeah so I guess like you know for for platform Engineers uh you have to use your own product right so if you're building an and you never trained a model before right or never deployed the model before uh it's going to be very hard for you to one empathize with you know your your customer and to like understand what you're supposed to be building right um so really dog food like the tools that that you're building right um from the data science perspective uh my advice would be involving your end uh and platform Partners as early in the process as possible um uh just to um I guess like validate the the approach and and really give them a heads up around like where you're you're thinking and like kind of the direction that uh you're moving in right so the the platform team needs to be you know six months to a year ahead of uh you know the business essentially the the product teams right um so really being able to to give that heads up is is beneficial and then they could also help guide you in terms of um you know different tradeoffs that you could make right there's a lot of different ways to to build a system um and if you don't have visibility into like what's what's difficult and what's easy uh you could wind up you know um incurring unnecessary costs right uh so from like a you know data science perspective if you maybe model the problem in a slightly different way then it makes it a lot easier on you know the platform to to support that use case right so really involve your your partners early on and having those discussions um really helps uh the the data scientists understand the limitations of the platform like what they can and can't do um and also uh helps inform the the engineering and and platform teams on what areas they they need to increase their investment in um and you know what what are the the pain points that um that EMES are having makes a lot of sense sense and appreciate all your Insight there um as we're talking about collaboration and and cultural um issues I am interested in actually a friend of mine um I don't know if you were were at Netflix when when he was there Eric Coulson who he was um VP of data science and engineering at Netflix and then um Chief algorithms officer at Stitch fix and I'm going to link to something he wrote recently for O'Reilly radar but essentially um it's a piece about that a lot of people hire data scientists for skills whereas we're probably better off hiring them for their ideas and not just have them as server centers right um so bring them into meetings where they can listen and figure out what the challenges people are having and I think he gives one really nice example of um data scientist in a meeting where product people and marketing people have done their like cohorted cluster analysis of users they've had users complete surveys and clustered people and the Clusters just have nothing to do with anything that's of interest to the business essentially right and a data scientist in that meeting said well I've got all the data now so why don't I go and do some clustering and figure out how we can use this to move the needle on product recommendations or whatever it is so hiring data scientists for their ideas more than treating them as a service center so I'm I'm wondering how you've seen this play out so in just ensuring that ml teams are seen as strategic Partners rather than just service centers yeah I mean I've seen the industry go back and forth on this where uh you have like a centralized uh data science org right and they almost act as like a Consultants within the company right um where like all right we have this problem that requires like some math or something like that and then you just like you know try to borrow a data scientist for like a quarter or something like that right and then you have like the embedded approach where um technically they have the same reporting structure so all um data scientists report to the same group but then they're embedded in specific product orgs right so like you know I saw this at Uber we had um you know uh uh data scientists specifically devoted for um for Uber Eats and then like you know some cohort of that would focus on pricing another one would be on like personalization recommendations um and you know then there's the third model where like okay they're actually just part of the team right and they're distributed and you know uh a manager could have software engineers and data scientists reporting to them right um I do see essentially like the a lot of problems with the centralized approach just because they don't have time to build up the domain knowledge um because a lot of it uh in order to be like a successful data scientists A lot of it is Data data wrangling right if you don't know where the data lives if you don't know what data is relevant to a specific domain even if you're like the best data scientist you're still going to struggle to solve the problem right because part of it's just like ramping up on on a specific domain that you're in right like how do you know what what metrics are important right like if you were to go in and try to predict uh I don't know like wheat yields um for for next year right like what would you even look at right if but then if you ask a um you know somebody who's like uh you know has like a degree in Agricultural Science or whatever um that has never worked in ml before they'll actually probably do a better job right uh give them like the you know Michelangelo UI from from Uber and they'll know what what features to pick and you know they'll have a very simple model but they would have picked the right features right so because they have that domain knowledge built up um and if you have that centralized org it's it's really hard for them to succeed um going through the to The Other Extreme right where you just actually have a data scientist as part of the team uh that works out but then you really need the manager to be uh to up level right um so it's very hard sometimes to appreciate uh if you you know if the manager has a software engineering background they might not understand the challenges and um the that data scientists face and might not be able to provide the growth opportunities right um so if you meet in the middle where you have like that embed uh embedding based approach or embedded based approach um that works out well um you know depending on uh the overall company culture right uh because you do have that opportunity to grow within the or uh at the same time you do have that ability to gain that domain expertise totally I um I love the idea of like even software engineering leaders or managers needing to upskill and also developing a shared language I've actually been chatting about this with a few people recently that um when talking about testing in llm or ml powered applications right if you're speaking with a software engineer they think a test needs to be needs to pass 100% of the time right A lot of the time um whereas we use tests in a way that if our tests are passing 100% of the time that that that doesn't smell right right in ml or or or or ai ai so maybe we should talk more about evaluations or teach s what we mean when we talk about Tes uh as well yeah yeah I mean especially for for traditional uh kind of software development um more determinism is expected right uh and that causes a I guess a lot of monitoring systems uh testing systems all that are buil or with with that in mind so when you apply it to ml um it doesn't work out as well right so like monitoring the performance of uh a deployed uh machine learning model you know it the you know traditional tools don't don't work as well right so you have like you know Prometheus for for metrics tracking uh uh which I think that I think that's spun out of uber if I'm think so I need to yeah like that doesn't really work for for ML systems right so you have like other other products that that you would use or build your own systems for it exactly um so we've got just a few minutes left I am wondering um the roles of ml Engineers data scientists software Engineers are blending and moving in a variety of different ways so and I think everybody is kind of wondering what what do I need to learn to to upskill now so I'm wondering what skills you think that will be most important for the next generation of ml practitioners and what people should be focusing on yeah yeah I mean it's it's a tricky question to answer because it really depends on what archetype you you want to follow right um so it's like a personal career uh decision uh some people really want to go deep in a particular area uh like you know if you want to do research you know that's a completely different ball game than if you want to just like Join one of the Fang companies and um you know you know still like work on ML application but in a more uh you know less researchy like capacity right um but you know if I were to give like general advice I would say focus on UNG gaining intuition um for uh going too deep in in any particular area right um you know I know like a lot of people when when they think of ml they're like all right let me learn like calculus and back propagation and and you know really go deep on on the math side of things because you know going you know going back to school a lot of the times is like all right like learn the fundamentals and you build upon that whereas for ML since the space is evolving so fast it's probably better to start with like all right don't worry so much about the specific math concepts understand the intuition of like how things work and you know the overall architecture uh of of certain systems uh and and ideas uh and then if you want to go deep in on the research side then like go deep on on the relevant uh um you know math or science that that you need to learn um but yeah I would say that the focus would be on on intuition rather than uh going too deep at least when you're when you're first starting out totally agree um on the same note I'm I'm interested well a slightly different variation on note for teams and individuals who are just starting to scale their machine learning infrastructure um what advice would you them on building reliable and scalable systems yeah I mean oh I guess like first off there there's never been a better time to to jump in than than right now just because of how far um open source has come um you know to the point where like like open source is outpacing Big Tech even in terms of like infrastructure that's available like sometimes it's even easier as an individual to like you know uh pull some like open source tool tools like metaflow and then run on your own AWS account compared to doing that internally at a large company right just because there's so much um Legacy software and uh specific checks and things that um make that difficult yeah um so like my you know main piece of advice there is use as much open source as you can um build in a way that uh you can evolve um uh because things will change um you know and you know if I could plug metaflow one more time um I would say like you know maybe use that as an example right so separating user code uh from like uh the code related to infrastructure and the related wrappers um allows you to really evolve your architecture very quickly um in like a plug-in playay type format building with uh uh really planning for deprecation almost of Civic systems right super cool man um so I'd like to just look forward for a bit we work in machine learning so doing some prediction isn't you know too unreasonable a request I'm wondering just where you see the biggest opportunities or challenges in ml infrastructure and applications over the next five years and I appreciate that's a very broad question but yeah yeah I mean it's it it's hard to answer and and you know I hate making predictions because I'm just like so wrong so often like just part of the fun though yeah yeah I mean like I remember uh I was one of the people that thought dl4j was GNA take off so like that was you know um a miss so like whoever's listening like you know you could maybe turn off the stream now because whatever I say is probably going to be wrong but but also like it was only a few years ago that Mark Zuckerberg made a big bet on the metaverse and then generative and like I love that move by the way and then he's like oh and just forgets about of course they're working a lot on all of that all of that amazing stuff but the way him and the whole team and Yan laon pivoted to all of this Transformer stuff like wild incredible stuff amazing amazing and yeah like you know maybe after the stream we could we could we could talk about some some specific stories there yeah um but yeah uh like in terms of like the the challenges for the next five years it's honestly just keeping up with the changes and being able to to Pivot rapidly right um you know again like the the space is evolving so quickly uh so much stuff stuff is happening in open source and it's it's not very clear what's noise and and you know what's substance right um so you know going to the second Point um it's it's also important to stay disciplined right you can't every Trend um but you do have to design your systems with like a added degree of flexibility of of flexibility in mind yep right um and I guess uh like another like maybe prediction in the next like one or two years I think the limitations of our current llms are are are finally going to set in um so there might be a period of like where we enter the the the trough of Despair right in terms of like yeah where people like understand that like actually you know what LMS can't solve everything right they're not going to place us all in the next six months um that you know maybe that that'll be like an over correction yeah uh so uh I do think that will happen sometime in the next couple years but you know if I could you know summarize in one sentence uh designed for change beautiful I um I think that's a wonderful note to end on I I'd love just to thank you for all of your expertise and and sharing it with everyone and I'd love to thank everyone who joined live as well it was a super fun chat um for us um and really appreciate you yeah yeah uh uh likewise and it's great to be on here and uh I guess since I am on here uh quick plug for for door Dash uh we are hiring specifically for uh related roles um so definitely reach out uh if you're interested in solving uh really open-ended uh uh questions in in ml designing uh building deploying ml applications uh or building the uh the infrastructure um to really enable that Innovation um we're we're definitely looking for for some great folks super cool and I'm actually just at careers. dash.com SLCC career area is the engineering page I presume is the the best one to share I I'm just going to share that in the chat now um definitely check it out everyone and check out the blog there's so much cool cool stuff there just to get a sense of what everyone's everyone's working on um and yeah it's really nice how on the careers page you share a lot of of what you're what you're up to um fantastic stuff well thanks once again uh for us really appreciate your time and um happy New Year everyone and looking forward to seeing you all up more fir side chats in the future and scene
Original Description
Join Hugo Bowne-Anderson and Ferras Hamad (Machine Learning Leader at DoorDash, formerly at Netflix, Meta, and Uber) for a fireside chat exploring the evolving landscape of machine learning and AI systems. Drawing from Ferras’s experience at some of the most innovative tech companies, this conversation will dive into the challenges and opportunities of building and scaling ML systems that bridge infrastructure and application layers.
Key Topics of Discussion
From Infrastructure to Business Value: Insights into how companies like Netflix, Meta, Uber, and DoorDash approach the ML lifecycle, from infrastructure design to application-level outcomes that drive business impact.
- The Convergence of ML Tools: A look at the trend of ML platforms converging to support both advanced users and non-specialists, addressing diverse personas and use cases.
- LLMs and In-Context Learning: How the rise of large language models is reshaping traditional ML systems, from tooling requirements to integration into production environments.
- Team Collaboration in ML Development: The importance of cross-functional relationships between data scientists, engineers, and platform teams to foster innovation and efficiency.
- Skill Sets for the Future: How the blending of roles like ML engineers, data scientists, and software engineers is creating demand for “full-stack” ML professionals.
- Operationalizing ML Across Industries: Lessons on scaling ML operations in sectors from streaming to delivery, with practical advice for companies at every stage of their data journey.
This session is ideal for software engineers, ML practitioners, and technical leaders seeking insights into the rapidly evolving ML and AI landscape. Whether you're tackling infrastructure challenges, deploying models at scale, or just starting with ML, you'll leave with valuable takeaways to guide your work.
00:00 Welcome and Introduction
00:10 Guest Introduction: Ferras Hamad
01:53 Metaflow and Its Impact
06:27 Diverse
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Playlist UU5h8Ji6Lm1RyAZopnCpDq7Q · Outerbounds · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Metaflow GUI for monitoring machine learning workflows
Outerbounds
Metaflow Cards [no sound]
Outerbounds
Fireside chat #1: How to Produce Sustainable Business Value with Machine Learning
Outerbounds
Fireside chat #2: MadeWithML.com -- Teaching Practical Machine Learning
Outerbounds
Metaflow on Kubernetes and Argo Workflows [no sound]
Outerbounds
Fireside chat #3: Reasonable Scale Machine Learning -- You're not Google and it's totally OK
Outerbounds
Metaflow Tags: Programmatic Tagging
Outerbounds
Metaflow Tags: Basic Tagging
Outerbounds
Metaflow Tags: Tags in CI/CD
Outerbounds
Metaflow Tags: Tags and Namespaces
Outerbounds
Metaflow Tags: Tags and Continuous Training
Outerbounds
Fireside chat #4: Machine Learning and User Experience -- Building ML Products for People
Outerbounds
Fireside Chat #5: Machine Learning + Infrastructure for Humans
Outerbounds
Metaflow Sandbox Demo: Free Data Science Infrastructure In the Browser
Outerbounds
Metaflow on Azure
Outerbounds
Fireside Chat #6: Operationalizing ML -- Patterns and Pain Points from MLOps Practitioners
Outerbounds
ML engineering vs traditional software engineering: similarities and differences
Outerbounds
Why data scientists love and hate notebooks: velocity and validation
Outerbounds
What even is a 10x ML engineer?
Outerbounds
The 4 main tasks in the production ML lifecycle
Outerbounds
Is the premise of data-centric AI flawed?
Outerbounds
The 3 factors that Determine the success of ML projects
Outerbounds
Fireside Chat #7: How to Build an Enterprise Machine Learning Platform from Scratch
Outerbounds
Run Metaflow on any cloud: Google Cloud, Azure, or AWS [no sound]
Outerbounds
Metaflow on GCP
Outerbounds
Fireside Chat #8: Navigating the Full Stack of Machine Learning
Outerbounds
How to Build a Full-Stack Recommender System
Outerbounds
Modernize your Airflow deployments with Metaflow - zero-cost migration [no sound]
Outerbounds
Easy Airflow DAGs for ML and data science with Metaflow [no sound]
Outerbounds
Fireside chat #9: Language Processing: From Prototype to Production
Outerbounds
How to build end-to-end recommender systems at reasonable scale
Outerbounds
Full-Stack Machine Learning with Metaflow on CoRise
Outerbounds
Natural Language Processing meets MLOps
Outerbounds
Fireside Chat #10: Large Language Models: Beyond Proofs of Concept
Outerbounds
What even are Large Language Models?
Outerbounds
How to get started with LLMs today
Outerbounds
LLMs in production
Outerbounds
Accessing secrets securely in Metaflow [no audio]
Outerbounds
Fireside Chat #11: The Open-Source Modern Data Stack
Outerbounds
Fireside chat #12: Kubernetes for Data Scientists
Outerbounds
Behind the Screen: How Amazon Prime Video ships RecSys models 4x faster
Outerbounds
Fireside chat #13: Supply Chain Security in Machine Learning
Outerbounds
Quick Delivery, Quicker ML: DeliveryHero's Metaflow Story
Outerbounds
Crafting General Intelligence: LLM Fine-tuning with Metaflow at Adept.ai
Outerbounds
Fuelling Decisions: How DTN Powers Gas Pricing and Data Science Collaboration
Outerbounds
From Kitchen to Doorstep: Optimizing Data Science Velocity at Deliveroo
Outerbounds
Building a GenAI Ready ML Platform with Metaflow at Autodesk
Outerbounds
Media Transcoding for 10 Million users and beyond with Metaflow at Epignosis
Outerbounds
Telematics with Metaflow: How Nirvana Insurance built a large-scale Risk Estimation platform
Outerbounds
Fireside chat #14: Generative AI and Machine Learning for Film, TV, and Gaming
Outerbounds
The Past, Present, and Future of Generative AI
Outerbounds
Building Production Systems with Generative AI, Machine Learning, and Data
Outerbounds
A Custom Fine-Tuned LLM in Action (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 5)
Outerbounds
Building Live Production Systems with RAG (LLMs & RAG: An Interactive Guided Tour Part 4)
Outerbounds
Better Relevancy with RAG (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 3)
Outerbounds
Working with OSS LLMs (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 2)
Outerbounds
Hitting OpenAI and Other Vendor APIs (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 1)
Outerbounds
Production Systems with Generative AI (LLMs, RAG, & Fine-Tuning: An Interactive Guided Tour Part 0)
Outerbounds
LLMs in Practice: A Guide to Recent Trends and Techniques
Outerbounds
Metaflow for distributed high-performance computing and large-scale AI training
Outerbounds
More on: LLM Foundations
View skill →Related Reads
Chapters (4)
Welcome and Introduction
0:10
Guest Introduction: Ferras Hamad
1:53
Metaflow and Its Impact
6:27
Diverse
🎓
Tutor Explanation
DeepCamp AI