Fireside chat #14: Generative AI and Machine Learning for Film, TV, and Gaming

Outerbounds · Beginner ·👁️ Computer Vision ·2y ago

Key Takeaways

This video discusses the intersection of AI, machine learning, and computer vision in the entertainment industry, with a focus on generative AI and multimodal systems for film, TV, and gaming.

Full Transcript

hi everyone it's Hugo bound Anderson here from out of bounds I am super excited to be here today to talk with uh ven misra about generative Ai and machine learning for film TV and gaming we're going to get started in a couple of minutes um if you could introduce yourself in the chat that would be absolutely fantastic let us know where you're watching from and and what your interest is whether you work in machine learning or generative AI um and and what uh line of work um you're in and what type of company you you work for we'll get started in a couple of minutes all right we have Mario from Peru hey Mario how are you she is great so we have someone from Peru of course venath is in um uh California and I'm in Australia so we're already um getting pretty global we have caros from Mexico working on Creative Education with AI we've got San Jose Boston Indianapolis ux researcher Gabrielle in Portland Haywood California Huntington Beach California Toronto Ivan from Sydney good day Ivan this is great we got people from all over the place all right we'll get started in one minute people um and if you could introduce yourself in the chat if you're just joining let us know where you're dialing in from what you're interest in machine learning and generative Ai and what what type of space you you work in that would be awesome as well so exciting we've already got um over 70 people here here live so thank you so much for joining wherever you're uh dialing in from and really look forward to um this conversation with vath all right I am too excited to not start now so venath why don't we turn our cameras and uh microphones on and get this proverbial show on the road good morning from Australia good afternoon to California good morning good morning how you good good yeah excited to talk about this fun fun topics for sure uh it is it is so exciting and everyone getting a lot of lot of love lot of love hearts and emojis and good times in in in the chat already and people from all over the globe um uh are tuning in we've actually we've got nearly 100 people already so that's super super exciting um no pressure at all just kidding of course um but um so I'm really excited to have you here today um to talk about uh all the exciting things that that you've been working on in particular gen and machine learning for film TV and gaming um before we jump in um I just want to say a bit about um metaflow an out of bounds um as as you all may know this is a fires side chat that's hosted by out of bounds and out of bounds we work on infrastructure and productivity tools for data scientists that allow them to focus on building models and doing science while having easy access to all the infrastructural layers such as compute orchestration and versioning um we do this through the open source metaflow and I'll actually um put a link to the GitHub repository in the YouTube chat um but we also have our products that support platform engineers and infrastructure engineers and um this type of stuff with that as well um and if you like what we talk about and like the the open source software please do star it on on GitHub um also if this is these types of things your jam um if you enjoy this conversation please do hit subscribe and and share with friends if you think it'll interest them um a bit more Shameless self-promotion but I'd have to fire myself if I didn't do such things um we have another fireside chat at the end of November which I which I'm linking to here um and this one is also going to be a lot of fun I think this is um with my my dear old friend Jeremy Howard of fast um AI on large language models for hackers so we're going to dive in and hopefully tell you all from Jeremy's perspective how you can get started with um llms uh straight away and that's one of the things I value about Jeremy is very very much skill-based and letting people know how to how to do what seems like magic and making it even more more magical um but speaking of magic um ven you've been um working in the machine learning space for uh a significant amount of time and so I'd like to just um introduce you very briefly and then maybe you can add a bit bit of flavor um but you're a machine learning scientist and software engineer at ROBLOX which is an online platform for gaming co- experience game creation that brings people together through Play Not only that though you used to um lead the artwork and video data science team at Netflix and that's how we got got introduced through my colleague and your former colleague uh vill who worked on metaflow at at Netflix um and there you used multimodal machine learning to assist creators um and analytics to inform creative decisions you've also worked at IBM Watson um and before that you're at Stanford and and MIT um on top of that I do want to talk about this a bit in a parallel world you worked as a technical consultant for HBO on Silicon Valley um including developing the fictitious middle out compression algorithm um so that's a brief introduction to to you but maybe you could add a bit more flavor to your your background um yeah no absolutely yeah I think you can definitely sense kind of a through line in some of those topics like I've generally been very fascinated by Creator experiences and the process of creation and how people create things um that's been a through line and and IBM also that was some of my work was was touching on that piece as well um the silken Valley thing was kind of a CO you know coincidental like just you know things just cross in the right time the right place but it also lines up with that same kind of story line and through line um yeah and I think like one thing I've noticed is that across a lot of these modalities there are some patterns around how creators operate how to work with creators and just directions of of how technology can impact creative workflows particularly ml based Technologies um yeah um happy to talk more yeah fantastic um and please anyone who's here um do ask questions in the chat and we'll do our best um to get to them the chat on YouTube um if there are questions you ask that we don't get to we're going to have an async asked me anything for a week or so on on slack so if you go to I've pasted our slack um uh workspace uh in in the chat so feel free to go there and ask any questions in ama guests as as well um so actually I I want to start Venice before before jumping directly into the topics today um I I want to tell you a joke okay so and you may recogniz um what do you call a murderer with moral fiber do you want me to throw the punch line back at you or yeah pleas yeah serial killer right a serial killer exactly um and so I wanted to open with that for two reasons one um one of the co-founders of out of bounds our CTO um Savin Goyle jokes that he's a Serial uh eating entrepreneur uh as opposed to a Serial entrepreneur but on on top of that um I wanted to bring that up because you gave several years ago um an interesting Ted Talk on the role of humor in in computation and how our relationships with computers can be funnier and friendlier um and i' love I mean that was before kind of the llm and gen Revolution right um but I think yeah and I I think it's actually very relevant to this conversation today how do you think about this now and why would we even want you could imagine a world in which computers do computation and we relegate that to them why would we even want computers to be funny and friendly and how can how can we manifest that yeah no it's it's it's and to be fair I don't think everyone necessarily wants that or needs it but I think there's there's value to be had from that and the the argument that you know you know that I I gave in that talk and that you know I I continue to kind of hold to is that just practically speaking we're spending so much of our times and our lives with these machines if you're you know honestly regardless of which field you're in these days and probably increasingly so in the future that's where a lot of your kind of mental energy is going that's where lot of your attention is going and conversely when you think about like the things that are most meaningful to you in in life These are you're often thinking about human experiences and human interactions and connections with other humans and and humor is is really a core part of that it's it's obviously not the only aspect of it but it's a big part of it um so I mean given that like I mean why wouldn't you want to bring more of that like deeper meaningful experience into something you're spending so much time doing um if you're you know if it if it works obviously that's a that's a big caveat and until recently that was a extraordinarily big caveat um still a big caveat um but you can imagine like you know instead of being a source of frustration could this stuff actually be a source of you know joy and and it is in certain cases already um but also naturally you just sort of people see people already kind of gravitating towards this sort of thing um so you know whether it's like you know cleaner UI like a great example is um it's a much more much Lovelier experience to chat with a llm that's proficient say Chachi BT or or Claud or whoever um versus going to an old school like search engine I think that's been relatively well established at this point it's a it's a much more human process in many ways and it it it it it it makes you feel much more fulfilled in the end and much less frustrated um actually going back uh to to web searches which you often have to do when the when for whatever reasons the the LMS don't quite cover use cases um it actually really underlines just how frustrating that that modality can be um and and how much you know actual value comes out from like the kind of I wouldn't call it human connection obviously it's not um but the kind of pseudo Proto human connection that you can get out of some of these systems yeah I love the idea of referring to it or thinking of it as kind of Proto or or or or pseudo human because these are the types of things we are trying to trying to create and trying to navigate I think um what type of human qualities we want them imbued with what human qualities we don't want them imbued with um and by them I mean these um models that we interact with generative AI machine learning databased databased models not datab based um although um Vector databased models um and I do think if we're talking about creativity humor seems in a lot of respects to be such a fundamental part of of creativity and and vice versa as as well um I also do think and this may be something we we get to but um humor itself is under Fire in a lot of ways currently um I think in in in in the Public Public space What comedians should be talking about what they they shouldn't um in terms of um being being sensitive to a lot of different groups and I think that's an important conversation to have but I do think perhaps products like chat GPT which let me get this right which um know one of the big things big wins of chat gbt is is its product right um like a lot of the technology a lot of the language model stuff exist pre predated cat gbt but the fact that it's such a beautiful product experience and that my aunt can can use it and interact with it um I do think one challenge with chat GPT is unless you specifically ask it to tell you a joke it is decidedly unfunny um yes and it Hedges a lot of different things has that been your experience as well yeah yeah no I mean and they've to their credit they're starting to play around a little bit with like allowing personalized prompts in there but as you know also they're not the only game in town like you you go over to some place like character Ai and there actually humor is like the core product offering in many ways in terms of the chat Bots that are being produced um so so yeah yes and and no y absolutely um which is almost always the correct answer right generally speaking yeah yeah yeah without thinking about something too hard that's usually where you want to start with yeah um and and so speaking about creativity and humor um jumping into the the real meat of this um when I started working for the first startup I worked for it was in the mid 2010s there was this show called silicon valon and I moved from academic research as you know to to Startup land and I I I watched the first season and I felt deeply seen um right so um I'd love just to hear and I'm sure a lot of people about your experience of working on like how it happened what it was like to work on the show particular with your interest in creativity as well yeah yeah no it was it was yeah I mean I mean these things are never really structured they just sort of happen somewhat organically so in this case um I I was working in data compression and information Theory his connections with machine learning at the time and graduate school uh my adviser is one of the well-known like experts in compression and silic and Valley and so uh you know HBO basically came knocking uh on his door like Curious it was just a producer named Jonathan Dotan who I'm good friends with though um and and reached out and just sort of was curious of hey we're making this show and compression is probably one of the central themes which was kind of a surprise why would you pick compression was sort of the in the early reaction um but then he he he knew I was like I was spending a lot of time in the drama department and and I had lots of interest in that space so he and he also figured that I would you know get along and be able to speak the same language as this individual which I was and and you know that's something we can talk about too which is really interesting in the creative space um but yeah I mean we H it off and and um at the time I don't think I realized how big of a deal the show is going to be it was just you know some HBO pilot that was shooting in you know they were they were shooting not in the Bay Area but they were they were scouting there and they were gonna set it in toen Valley Etc um but I don't think until we saw like some of the first drafts and we saw like the ways they were using the technical like content and the depth to which they were reaching for authenticity there um that it became clear this is sort of a different creature from like what you've you know typically seen um yeah a couple interesting things I would maybe out there one maybe along those lines I think my theory of like why this show resonates so much A lot of it obviously the writing is spectacular the performers are great but I do think some of the heart of it is actually um Mike Judge himself is actually just a deeply deeply technical individual like he gets it he has a level of curiosity about this stuff that you don't typically see from show showrunners who are operating in the technical space um I mean he used to be an engineer in in Silicon Valley uh you know prior to going to Austin and doing his whole thing with us and Butthead um but yeah he you know for I can give you a couple examples here like um I showed up for instance one of my first like pitches for them was I I made a little deck of like some um they they had some ask I can't even remember at this point what it was for um but it was a breakdown of some algorithmic basis for some conversation um and I had kind of like bsed a little bit here kind of fudged some of the edges to make it look nicer and their feedback was actually no this is actually not legit enough we actually need this to work and it was it felt like almost like a pure review in terms of like the level of depth that they were going into Mike in particular um another one is like in the first episode or two there's some discussion around doing data compression or doing um processing inside a compressed data space um which to me was a really provocative technical research topic and something you can imagine Building Systems around or innovating around and I had assume that my adviser saki or someone had like given this cue but it turns out that was actually like coming from Mike himself and he had like a strong inclination there um one last data point I'll share there also um um I was always curious like why are they focused so much on lless and not lossi and it turns out Mike actually had the same bot like his initial um impulse was this should be about lassi compression because there's so much more that we can squeeze out of that which has proven out post like you know neural Nets and everything um but the the the the the trade-off was lossy sounds bad to an audience right if you don't have the full context lless sounds way better uh which is which is funny but but um but just like level of of of aess of like some of these nuances which you wouldn't typically see and and the engagement was much more of like a conversation it wasn't like a transactional fill out these charts kind of thing yeah it a lot of fun it was a lot of fun yeah amazing and thank you for that Insight um I also think ly may sound a bit worse to VCS as well but that's probably that's a fair point yeah yeah General another day um before we move we just got a we've got a lot of comments about this but um someone's written LOL that show is amazing Long Live middle out and I think that's yeah that's so I still remember the first because it was actually there's a funny story with midle out where um initially um jonan had just sort of described and then we had just gotten sort of a rough idea of like almost like a speck of the the parameters this algorithm needed to fit into and we didn't really have this whole story context for it yet um this is they were still working on drafts and such um so we it was really weird like ask like hey we the the algorithm needs to fure a back and forth motion and it was just like some really odd technical constraints um and then when we finally got the first draft and and we were looking at like why that it needed all of this it was like oh man this is um next level um yeah definitely like jaw-dropping moments in there um so I'd love to now jump in um and and start off by uh talking about your your work at at Netflix um so at Netflix you talked with um and worked with all types of creators and creatives um what happens at a place like Netflix when machine learning Engineers work with content creators such as Direct editors and 3D artists to name a few yeah yeah no I mean I think it's I mean first off I think like you know there's there's a a failure mode here and there's like success modes and there's multiple failure modes right when you're having like two very different like folks with completely different backgrounds incredibly different diverse backgrounds and and and understandings of what's important and not are highly diverse across these two spaces um one one fill your mode is you know basically you have like this Chasm almost in verbiage and like Chasm of like assessment of where the opportunities are like if you typical machine learning engineer who doesn't necessarily like embed themselves in the creative world may have certain assumptions around what the pain points of a Creator are and and might anchor on them they might have trouble even describing like what the opportunities are conversely like you have creators who may not appreciate what's feasible or or what's not feasible um there's a old XKCD I think uh where where you know talks about like it's incredibly difficult sometimes to explain to people the difference between something that'll take like you know 20 minutes at a computer and something that'll take five years in a research team of like 20 to solve you know um and and and that's sort of true so I think like one failure mode is where you just have like almost like failure to to engage right and you have people who are just not not connecting on on the actual problems and a lack of curiosity and openness to understand what the real problems are um which are almost never what you expect them to be um for instance like if you're if you're working in like trailer helping trailer editors um your assumption might be okay let's automate trailers trailers are just summaries they're like video summaries of a movie so let's look at text summarization and then you know the moment you start actually engaging on the creative you realize that's not at all an effective uh summary of what a what a trailer is it's a much more complex creature it's totally different from a summary it's its own story practically and and and the pain points that you would imagine are are not what you would expect um example here um you know you may think of like hey we we need to find the really Snappy moments and like pull those for somebody or or or figure out how to like tell a story but honestly like if you're even just a make allowing someone to just search through the text of of of a subtitles of stuff that's already a huge win you know so it's like there's often like these logistical pieces that um are maybe not so sexy that are often like a source a lot of value um and and conversely you know like there's also a lot more ambitious stuff you could be doing on the tooling side that um sometimes you know like our creative partners aren't even aware of like that are in the realm of possibility um the other failure mode I think that's worth calling out um is what I would call like more of a transactional kind of model where it's almost like differential where you have like ml Engineers who are basically just asking their creative Partners like what should I solve just tell me what to do and I'll go out and I'll do it and you know it's it's almost like um it's a deceptive failure mode because it doesn't feel like a failure mode because it feels like you're you're solving actual problems for actual people you're airing on the side of listening very carefully and you're you're you're you're focusing on the customer and their needs but I think what it misses is the ask for a really wants B kind of scenario where it really requires like engagement and pushing and partnership and prodding and and in some ways both of these sides need to be able to really engage with the other like on the creative side we really needed these creative strategists these folks who are able to to really put on like their technical hats and speculate about like what was feasable with the technology and really push the technologist to like think about hey could we actually do that and convers we need a technologists who are recognizing hey there's this like you know cool technique that I think we might be able to use is there I have a feeling it might be useful for this like maybe this match cutting thing or something what do you guys think um and and without that it becomes like honestly just a very low ceiling of of impact um anyways I've been talking for a while but does that does that make sense yeah so that makes perfect sense and it's really exciting work and I'm interested in what type because we have a technical audience here what type of tools and techniques you were using at the time because you could imagine that for some of this maybe logistic regression is fine but maybe you need to do like some sophisticated deep learning and like significant data engineering and that type of stuff as well methods and and software did you use absolutely and and I'll caveat this also with like I I left Netflix like almost three years ago at this point and that and that team is awesome and they've been doing amazing stuff far far more impressive stuff now than what I what they were doing when I was working with them um and they're good friends and I I yeah I would definitely call caveat anything I say with that absolutely and obviously a lot has happened in the space since then to say it defeated mildly um but yeah I mean I think like I'll I'll call it a couple of things one um what you often find with a lot of these scenarios where you're trying to help creatives or help honestly in most spaces um there's a bias towards generation that always sounds really exciting and sexy and like it's it's it's machines Being Human but usually where the vast majority of your initial value comes from is from retrieval it's it's not from um from going out and creating new stuff it's actually making better use of the content that's already out there um and this is particularly true I would say in places with like really premium content where you know the the you know the bar for a Netflix show is like up here you know compared to um you know a Tik Tock video that I'm you know shooting with my kids or something you know I can I have a much lower tolerance for for shipping quality product there right um so retrieval is one thing and retrieval comes hand inand with there's really two pieces to that um one is as you mentioned like the data engineering side of this actually having access to this data the ability to index it the ability to retrieve it at scale U things that you know I would accurately call infrastructure Investments That many places frankly haven't made enough of and it's a bottleneck because without that you can't even do retrieval let alone uh you know retrieval with characterization of your content let alone training generative models like all of that is sort of off the table um the next layer beyond that is and kind of alluded to this this just there is is characterization right like it's it's you maybe have this massive catalog of you know whatever content that you know your company or your creative organization has been uring over time um but if you don't know what's inside those black boxes those black boxes are useless right they're only useful in so far as they're characterized um and characterization interestingly it's almost like a discriminative problem compared to the generative problem of creating stuff right so you rather than creating a movie it becomes a question of going through all these clips and characterizing everything that's happening in them in a way that allows you to retrieve them as need be in the future um so you know a lot of the initial value um when you're digging into one of these spaces is often from those two pieces of characterizing better and building the Machinery to retrieve it properly um and then the last thing I'll call out there is once you have those two pieces in place even before you touch all the geni train and all of this that itself unlocks a whole ton of use cases because you almost have this like nice little sandbox playground where you have all this data you have this nicely labeled and you have all these like very long tale of creative use cases um which is another thing that's really interesting about creation is that it's not like um I don't know I don't want to demean any particular profession but it's not like a profession where you're you know repeatedly doing the same task over and over it's it's an extremely thick tail of the kinds of tasks and needs that you have um which is part of the reason why we think of it as such a human operation um and and because of that you know like there's always going to be you know a long tale of ways people want to use this data and repurpose it um whether it's like you know stitching together stuff that looks kind of similar or here's a story let's match stuff that might fit the story or you know even just like I said just retrieving like a specific line of dialogue and the scene where it is uttered um there's there's a very thick tail there um so again that's that's pregeneration so this is all just like characterization work um and and I would say that was where the bulk of our effort was um characterization and just to be clear characterization is not a trivial effort um like you're talking like large scale video models that are able to you know comprehend um this multimodal creature called you know streaming video which is you know there's text in there there's images there's video there's audio um there's a lot going on there's metadata on top of that there's user data there's there's a bunch of Dimensions to this stuff um and it's um and it requires like some serious investment to be able to actually characterize that effectively um but again things have shifted and a lot of those problems that were you know very heavy lifting you know three years ago are probably significantly you know more tractable today and the tools are probably significantly more commoditized and and democratized so I know you're not you're not there anymore and I do want to get on to thinking about generative AI in your work currently but I'm wondering um you know how else do you think um all the work you did at Netflix would change now um given all the new tools and capabili yeah yeah no absolutely I mean I think like there's there's a couple really big changes that have happened with Gen right like and and the progression of the ml space in general um so one is like just really obviously the quality of generation content has just skyrocketed right we've we've seen an kind of like an an explosion in quality um particularly obviously we've seen a lot in the image domain I think video is is coming it's not here yet um you know folks said Runway may have another perspective on that but it's in my opinion it's not quite here yet but we're getting there um 3D stuff obviously we're we're deeply invested in that space here at ROBLOX um but but the quality improvement that what that happens to do is it just it it's really like this is like a combinatorial like playground again going to that analogy of you have all this characterized content and now you also have models that can leverage this stuff right and and the more you know models you have that are able to produce like higher quality stuff the more degrees of freedom you have to play with right so you know like the kinds of things that just weren't possible before like previously you would need to rely on heavy human effort and slow iterations that's the more important part is the pace of iterations you would need to rely on that for certain operations say it's generating a proof of concept like image right or or condition generation of something like these are all things that um previously just weren't in the scope of of of automating and so the the speed and scale of what you could do was was severely limited so so quality is one piece um the other piece I'd call out around ML and and as its penet Creed more workflows I don't think this one is appreciated as much or called out as much at least um it's what I would call like the velocity of experimentation and the velocity of being able to try out new techniques um so you know an easy comparison point is Imagine like you were an image processing or or video processing engineer in like the early 2000s right um ml was certainly a thing at the time but the quality bar wasn't at the level where you could actually really leverage it for a lot of the activities that were aaging it for today so but that doesn't mean you couldn't do a lot of the things that we're trying to do today with algorithms like you could you could green screen right you could you could you could do uh you know there's there's all sorts of things you could be doing you know compositing Green Screen there's a long list VX effects um you could do a lot of that but the problem is it takes extreme amounts of domain expertise and ex extreme amounts of like human fiddling and iterations and algorithmic Hands-On development to to get to a reasonable level of quality there and ultimately in the end what you what you have is is finely tuned it's it's involves a lot of effort and it's it's still probably you know somewhat pass it's not quite at the level that it is today and so what that ultimately means is that it it takes a lot of effort to spin up a new tool like every tool is kind of its own little like island of of engineering effort and and prototyping that is highly non-trivial um what ml has done in my opinion one thing that isn't heralded enough is it's really lowered the bar um and and increased The Leverage that you get out of a single individual which these teams invariably are is like just a few individuals working together in prototyping stuff um because all you really need is data at scale you need compute and the hopefully you're you're you're operating near a state-of-the-art that's you know reasonable in quality and at that point you know it doesn't really matter a whole lot whether you're working in the textual domain or the video domain or the music domain or um let alone within the video domain you know like editing in this particular manner versus this other particular manner um there's a lot of of fluidity and ability to kind of like rapidly iterate and prototype and build new algorithms as long as you have the data to support it um so that that's another thing which I feel like has been under kind of appreciated or or or under woried around the benefits of ml um it's not just about quality yeah absolutely I appreciate that now I I do want to jump in and hear about what you're up to at at at ROBLOX and we actually have an interesting question which I want to move towards J J Perez has asked are there any notable challenges or limitations when using generative AI in gaming uh such as in in in in Roblox and I want to move towards that but before that i' just like to get a general picture um of how you think about using generative AI at ROBLOX and in particular um I know that you're particularly interested as we've already discussed in multimodal gen systems so maybe you can just give us a bit of background from your perspective on multimodal systems where they came from where they're going um with respect to the linear content creation and experiential systems you create at roblo absolutely yeah and even just let me comment on domain also like I think part of the reason why multimodal is so exciting and and in the gaming space particular is I think games are probably the most multimodal kind of like creature on the planet like you you you have just basically think of a modality and it probably exists in the context of a game right you have music you have speech you have 3D graphic I mean 3D meshes you have textures around those meshes you have animations of those meshes you have 3D worlds that are getting created arranging these things around you have code that's being written that's defining how these things behave you have highle game design which is like another like kind of layer above like all of this stuff um that and there's user Behavior layered on top of all of that right so it's it's an insanely multimodal kind of space um which is why I'm actually kind of bullish and I'll I'll explain why about the future of of gaming and its intersection with ML um and the reason is also part of the reason why people are so excited about multimodal stuff in the ml World um which is what we found you know generally the general principle in ml that's sort of emerged in the last like 10 years or so is that um wide and Broad beats narrow and deep most of the times and and what I mean by that is if you're training a massive model across a ton of different tasks across a ton of different data sets throwing a bunch of compute at it it'll eventually outperform any narrowly scoped model mod on any one of those tasks and and often times dramatically so um so that's something we've seen in the context of you know obviously llms which are fundamentally like multitask and and infinite task in some ways right um and and multimodality is basically like the next layer the next dimension of that expansion right where you know once you once you run out of text right where do you go you have all these other modalities that are also encoding aspects of the world and about reasoning and about behavior and tasks and such um so it's it's naturally like the you know putting on your futurist task for a second um the the direction you'd expect like continued Innovation and and benefits to come from these models is as they absorb more modalities um and and in my opinion again gaming is one of those places that's really primed to profit from this because of the sheer number of modalities that it touches um sort of like a convergence of media kind of moment is is is how you might think about it you know um but yeah I mean so I think like the other thing about multimodality aside from just the raw Improvement of the models um goes back to my earlier Point around combinatorics um and what you find is anytime you throw an additional modality into how a model works or how well it operates um you you sort of are are kind of adding another dimension for the kind of applications you can you can explore and and and be capable of of of supporting right so um for instance you know lmms are great if you're working with Text data and working in a text domain they're awesome and there's a obviously a super thick taale of use cases people are discovering around us right like everything from writing marketing copy to making decisions about you know how to service a user um all sorts of things right um but ultimately like when you start throwing in say even something as straightforward is just like images like wiring in images like gp24 V is a great example of this um just anecdotally in the industrial sphere I don't think people even just judging for my own network I don't think people really realized ahead of time just how many use cases this was going to surface for them and we're starting to see this kind of in real time like and and the reality is that pretty much any data source any any you know how do we experience the world it's multimodal so like there's so many data has so many artifacts that are fundamentally Crossing text and video um the ability to apply you know what you're doing in what you're learning in the text domain like the capabilities of llms to the visual domain naturally opens up a bunch of like applications and opportunities um so anyways yeah maybe to summarize I would say um models getting better right that's that's that's one thing we care about with multimodality um I think it aligns with some of the data sets that I'm particularly curious about and interested in like gaming right and and and lastly um this point around just sort of like the combinatorics of the use cases of supports once you start throwing in additional modalities um also I think it's worth calling out like we're very you asked about how things have changed I would argue we're very very early on this multimodal joural journey um like we're just scratching this surface I think like there's um and and and what we're seeing is like dramatic advancements that are being made you know in some like logarithmic time scale as so um so anyways I think we're going to see a lot more from this in the kind of like years and months to come um yeah sorry I mon for like no that was great so I'm interested now in in diving into what what are the the most exciting use cases of gen and multim modality for you at ROBLOX currently yeah yeah I mean I think like so ultimately like when you think about characterizing what's going on inside like a 3D space right this is just one example I mean there's there's a lot of things and this is true for I think any game developer any game system out there like like there's the understanding what's happening in this is super ambiguous right like when you're when you're looking at let's say let's say you're collecting you have a mobile game somewhere and you're collecting this your logging data around what your users are doing that's giving you actually a pretty limited view into what they're actually doing like maybe you're logging something about the Milestones they're clearing but you know how much are you seeing about their experience in this 3D World or or or what they're seeking to do or the circumstances of it um and and being able again going back to my point around characterization that data set is is really a black box for for most folks in the games industry today right um knowing what's happening inside the games and what's happening with their player base and what they're doing what's it's basically a complete blackbox um but you know if if you were able to start like anchoring on some of this multimodal stuff and actually be able to extract out like context whether it's from screenshots or videos or text shat or maybe combination of all these sources of evidence um it allows you potentially a much richer source of characterization of your game world right um and and that you know as I mentioned like I'm always interested in the combinatorics and and once you have that kind of like data set it empowers a lot of different things right whether it's like helping people B make make Superior games whether it's accelerating the velocity giving them insights on what's likely to be a better game model there's there's a whole number of Dimensions this can this can go in um but you know like there's the standard saying of of ml which is like garbage and garbage out if you if you don't have the ability to actually see what's happening in these games or see what these games actually are there's very little that you can say about them right yeah um I'd also argue like so there there's two use cases frankly that I'm really excited about that's one is like in the gaming space and and understanding what's happening in these 3D worlds um with a level of telemetry that honestly doesn't exist in the real world right like we we I don't live in a house that's like covered with cameras that like you know is is recording everything I'm doing and and all that in a certain anonymized way we we don't live in that world and we probably shouldn't um but um but but we do perhaps in the in some of these gaming spaces and that's a really provocative concept um from a perspective of building things and and learning how to make better things um the other use case that I find really compelling um is more going from the virtual world context to the real world context um I do think like you know this this might end up being sort of the Saving Grace of a lot of the investment in AR and and and and you know reality labs and their camera and their camera glasses or whatever else is being built in that space um because ultimately like that is the best the way we experience the world is multimodal and you know ideally if you have models that are able to comprehend at that scale the the scope of applications is really unbounded at that point um and yeah I think even in the recent like earnings call like Zuckerberg basically called out it wasn't earning call sorry it was I think it was meta connect um was calling out literally that fact that like hey this has actually kind of surprised us that we're kind of positioned very well for this kind of stuff and and I do think that's a really compelling use case but maybe a a a special case of a broader thing which is multimodality aligns with human experience and therefore it aligns with human driven data sets as well I love it um I we do have a set of questions which I'm trying to kind of like massage together in some was it has to do with the the experience of you don't have window open to you Hugo suiz for you I really should have um yeah I mean a plugin to get this chat get the YouTube chat into chat GPT uh immediately and do a summarization would is a nice product so um uh Jesus has asked what what role does machine learning play in player prediction player Behavior prediction and dynamic in-game adjustments for a better gaming experience um uh moan has also oh no I'm sorry that's another question I I want to get to we have a question around um if generative Ai and machine learning can help us um with NPC character behavior and prediction and and World building so I I suppose a general way to kind of summarize these questions is how um these types of tools and techniques can impact um The Experience within the game yeah know absolutely it's it's a great point so I'd call it a couple things one some of these things are are not necessarily new things that are enabled by gen so so for instance the point around predictive behavior in game like like we've been able to do this for a long time like we are offline predictions and stuff and and chances are if you participate in any like reasonably well-resourced like free-to-play um MMO or or or live Ops oriented game chances are there is a fair bit of predictive Behavior happening inside that game um whether it's like what particular like powerups you're PR presented with what monetization angles they're taking advantage of um there's there's chances are there's there's a lot of ml behind the scenes that that's already there um I do think though that perhaps like to to underline the question I think like this does expand the scope of what's possible right for sure in terms of contextual awareness in terms of being able to kind of Riff and create new Concepts on the Fly there's there's a lot of things that potentially become enabled by this um I'll I'll call out a couple examples so um I think like you can imagine um today and a lot of the stuff comes down to robustness and and quality and robustness and that kind of tra off like in a world where you're able to generate content dynamically which we're not quite at yet I mean you can see obviously Twitter videos of people generating like 3D meshes that are textured and everything but if you look at like the time it takes to create those and the reliability of getting good outputs where we're maybe not quite there but we're close um and we have some interesting stuff happening in this space coming out very soon for the record um you can imagine that starts enabling like a lot of very novel kind of like gameplay Loops where you're actually not just serving pre-made assets but you're actually potentially doing things dynamically in game um and that's a really provocative Direction I don't think we're quite there with the tech yet but the the momentum is definitely heading in that direction and it'll be interesting I think the thing I would call out here as with the NPC question is um one lesson we've learned at ROBLOX um and that I've learned working with creators over the years is that your imagination as an engineer or even as a single Creator is woefully undermatched compared to the imagination and creativity of a Creator community and the things they will do with like some of these Concepts um will is is far beyond what you can probably imagine or or even think to to experiment with um so I think some of this is actually we're going to discover in the next few years what the real use cases are um so like with the content generation the ways that people can start exploiting that we don't really know yet um but but we're GNA find out um NPC is another example where there's obviously a tempting kind of angle there around um you know like hey what if I could just go and talk to this NPC and then that'll be like you know it's a random person in this game and I'm able to have an actual conversation with them and maybe the quest line like dynamically adapts to that like there's all these Concepts there and I would call that like having talk to folks who are poking around some of these ideas and and and experimenting with them I would say like the the real value has yet to be really identified and unlocked there that's my personal again take not not speaking for Roblox or for Roblox creators um the because ultimately like there's traditional games are designed are not designed for that right for NPCs to go off on like random non seiters and quests to change dynamically I think what we'll probably see is is some new perhaps genres or mashups of genres that are designed around this mechanism and doing interesting things with it more so than just like you know your next time you play Call of Duty you can like talk to the person before you you know shoot them yeah you know it's a it's it's likely to be something different from that amazing so as someone who uses I mean a lot of what we've seen in the gen space um up to this point have been proof of Concepts um and you're someone who uses it at at work so I'm I'm wondering um what are the most critical technical challenges that that you see actually using it at work on a daily basis currently yeah yeah and i' say even I would distinguish between using it as in like obviously I go to chat GPT and asking questions but that I don't think that's what we're talking about we're talking about building these systems Andy to deploy them and build products around them yeah so yeah in question I really love this one um yeah I don't think you know like over the years there's been some like Notions of like best practices that have kind of emerged in the ml space it's undermined a lot of historical software best practices and also reinforced it in other ways and I think we're we're early on that same Journey with geni um and there's some things that people just don't talk about a whole lot um and some of these things are kind of obvious when you talk about them but they just don't get enough air time um so one example of this is um you know if you one of the Hallmarks of a good ml engineer is focus on evaluation not focus on models like people associate MLS with models but it's not a modeling space it's an evaluation domain um and and with geni I think we're still figuring out what the heck to do with the vows um and the reason is this right historically with a more discriminative kind of case whether you're retrieving stuff for people or you're you know classifying is this a dog or a cat or whatever like kind of digestive use case you have for for ML discriminative digestive use case um it's very easy to come up with evaluative metrics that align with user experience right like how often are you getting the right animal or some metrics that are derived from that or or there's plenty of ranking metrics that are floating around one can one can reliably kind of use as a proxy for Quality um with generation though and this crosses across every modality um the challenge is your your something new is being created and and you have to evaluate the quality of this thing that's new and this has been a thorn in the side of generative work honestly for like the last like 10 years 10 plus years it was a problem as soon as Gans showed up it was a problem before Gans and it's still a problem today um and the reason is you you you don't have ground truth right this is something that's by definition different from your ground truth and and we've developed some you know mechanisms and some you know um things that are that are quantifying that um but I think it's still very much of a moving Target um I can share some of the the the kind of rules of thumb and some of the tricks we found that have been effective in the space um but generally speaking I would say the meta rule that's really important is to spend like probably 80% of your time UPF front just aligning on what your evaluat strategy will be and and only then can you actually start building the system and iterating on it and improving it um otherwise you very quickly end up in this really amorphous uncomfortable space for Gen where you're you you have a new model version or you've changed something and every time you hit the button you get something different and you have sometimes it's better sometimes it's worse and you're not really sure maybe you have a bunch of humans who are evaluating it but now you're bottlenecked by having hundreds of human reviewers or hundreds of human reviews before you can even assess the quality of your latest iteration um so anyways automated Deval super super important underappreciated um and happy to talk more about some tricks that people have kind of discovered in the space some so there's some interesting threads here um but still early days I would say is the is the bigger takeaway yeah and something we do see a lot is you know people evaluating by saying you know it passes passes A vibe check for for example yeah yeah and yeah and there's there's a lot of like fitting to noise in that case and oftentimes what I found like a good test of this is like once you have reliable evaluative metrics go back and and see how many of your Vibe checks were actually real and fake and it's it's it's a depressing kind of stat um I think there's also like this phenomenon where we're all very excited about the space and anytime you're you have a new feature or a new thing that you're doing you're very there's a big confirmation bias in there you want to see the good in it you want to see it as an improvement no matter how disciplined you are you really want to see that um and and oftentimes it will lead you astray um and and the comeuppence doesn't come until you start putting it in front of users and you realize this is actually the same as before or worse yeah so then how do you actually think about we've kind of touched on this a bit but very explicitly how do you think about moving gen from proof of concept to production um and what is actually working currently yeah yeah so I think like one thing that we're finding there's basically like two general approaches that I've seen be somewhat success or successful I think evals is like 90% of it frankly I think set up your eval system so that you're you're getting authentic feedback that's reflective of your users you have evaluative prompts that reflect evaluative prompts or whatever input there is that reflects your user needs and metrics of evaluation that line up with user value and it's almost like a paperclip maximizer if you set that up properly and you have like a team of Engineers who are competent and motivated they will move that metric in the right direction and that that I've seen happen over and over again which shocking rapidness it's it's it's quite striking how quickly once you have that system in place things just like spin forward um but I think like in terms of developing those things um a couple things which I've observed uh one approach is you know you actually go the model training route I this is kind of inspired a little bit by rhf but basically you you look at you basically explicitly are training models on sort of these unit test like prompts um so perhaps say you have an image generator you have certain prompts I'm just using this hypothetically you have a certain set of prompts you're evaluating on you have a bunch of labeled outputs from your model um or other models or just images in the wild that are good and bad you train a model specifically on for that prompt that's telling you what's a good output and what's a bad output um and then you apply that as your scoring mechanism right and it's noisy it's not perfect but it's it's surprisingly effective um uh particularly I mean depending on like how many labels you have and all those sort of usual parameters um but it's sort of like evaluating via model approach that um does actually unblock you quite a bit um the other one which is even more I think relevant for a lot of these agent like systems and things that take actions um are explicit unit tests um so you you've seen this a little bit with like code completion like and and and code generation where you can literally write unit tests like it's a it's a function you want to fill out and you've written a unit test and you're checking whether it passes that unit test but you can apply the same principle to other modalities so um say you're you're creating a um I don't know and and and um I'm trying to think of a good example you're you're you're generating like say a scene of some sort and a common failure mode is that things are overlapping in your scene so you write an overlap detector right and that's your unit test um and and you now have something that's you know deterministic but it's giving you a meaningful signal on whether the model is getting better um you have a kind of library of those you're actually starting to be in a pretty good place in terms of detecting progress um yeah so anyways um those are the two big levers I found but um neither of them are sort of easy they require like much more effort than your traditional evaluative of like systems absolutely and I think evaluation as we've been talking about is is absolutely key something we hinted at earlier is at the other end of the pipeline you know you mentioned garbage in Garb garbage out making sure your data is is yeah what what you think it is and and labeled well and all of these types of things so I suppose having a maybe you could talk a bit about the importance of having good data and actually how the skill set required here is not dissimilar to kind of the classic skill set of a of a data scientist right yeah yeah it's it's a great Point yeah I think like um ml has basically converged it's like a bifurcation of two disciplines one is like the data prep and the other one is like the the system building like the engineering side um so I totally agree with that characterization of the data piece being like absolutely critical here um one there have been a couple interesting Trends there um one is you know you hear about these massive massive data sets right and these efforts to like create the data set the web craw scale data sets that you need to train an llm for instance or the web crawl scale data set you need to train stable diffusion or whatever right but in practice what I've generally found the trend has actually been that the the trend in practice has been towards smaller data sets that are higher quality being more impactful um so as these Foundation models get larger and larger most of the people on this call are not going to be training a foundation while they're going to be using it and the kind of data that's really valuable for them is actually doesn't need to be very large at all either you're fine-tuning or you're few shotting even more extreme case um but but when you have a 100 examples the goal is really to get the best possible 100 examples so that your your fine-tuning or few shotting is like really perfect um so it's actually I feel like the trend has been in many ways away from data scale at the application layer and more towards data quality um which is you know a little bit different from you know if you rewind the clock 10 12 years and it was all around like you know doing your massive math produce jobs and data mining to get your data sets or training your models but the story is a little bit different in 2023 absolutely um so we'll have to wrap up soon we do have um so many interesting questions I'll get to a couple of them but um the ones we don't get to please do join um because we have have a lot um please do join our slack which I've linked to in the chat and the channel is AMA guests there um Mustafa has an interesting question um which I think plays into the deployment story as well which is how to machine learning algorithms assist with personalized content recommendations for viewers or players um so we're talking about recommendation systems here right which I think with your experience at Netflix and how you what you do at ROBLOX of course is very very interesting but I I do want to preface it by saying one reason this plays into the deployment story um and business value story is you can build the most beautiful Ensemble model that has you know significant lift over all models um but if if it's too expensive to put into production or takes too long or yeah latency issues the these types of things it may not actually deliver the value you think it is and I think absolutely the typical example that I cite here is the final model that won the Netflix competition right there were precursors to it which were used in in production but the final I think it was belore team or something like that I can't remember the name but yeah it was like this yeah um it was a model which was a highly performing model but wasn't put in production in in the end because of other other costs but to be clear I think the model also now this is actually really important when think about business value the model was recommending to people pre- streaming as well via however Netflix did it then so then the business pivot to the streaming model and of course Netflix has finally deprecated sending actual yeah what a yeah I mean I was actually surprised that I I thought that would have happened ages ago but maybe that's just just me as well that was one of the last holdouts by the way I was still a subscriber to the DVDs until yeah last yeah um so all of that is is the preface that this type of stuff doesn't exist good models don't exist in in in a vacuum right um so yeah how how do you see ML and generative AI helping with personalized recommendation both plac and Roblox yeah so that's that's a huge there's a lot to boil there so like I think like a couple thoughts um one I think like on the point of like latency and such like I it's hard to underline just how important this is like people in theory even like product managers will tell you like you know let's focus on quality let's not worry about latency we want to deliver a quality product first I think the truth is latency is a huge factor in quality there's a there's a great little block post around like some of the folks building like the GitHub co-pilot like original version and and it's one of the big takeaways from like the stuff they said and shared is just the importance of like low latency almost at the expense of everything and the sheer value delivery users based on that short turnaround time um so it's absolutely true I think like tight slas are a big part of this with geni I mean that that puts it in in Stark relief because that's the fundamental challenge with a lot of this stuff is how do you get it out quickly and I think we're still figuring out the mechanisms so you know there's there's tricks in this space you know these models are getting faster and and there's tricks they're playing in terms of how to do inference more quickly there's also tricks you can play on the prompting side around trying to like Minify your outputs because the output tokens are where you really get get hit hard on the latency side um input tokens you can actually scale much more effectively um uh trading off Which models you're using depending on like their complexity of the query there's like there's like a whole host of tricks you can kind of play here um but ultimately I don't think we're quite in a happy place yet on the latency side and that's going to be a paino I think for some time um so just plus one on that um recommendations versus generative Ai and personalization so I would say generally recommendations and the big lesson from that is user data trumps everything when you have it um and that's that's the critical caveat and I think in cases where you don't have it that's the cases where I think gen has a lot to offer in addition to like the conversational modality of interaction and all of that um so you know if you're if you're if you don't have a lot of users yet and you want want to be able to recommend your content catalog to them what better way than a Content based recommender powered by J right like that like LM um similarly if you if you have new content that you don't really know how users are going to engage with what better way than something that's characterizing it and doing a Content based recommendation on top of it um so I'd say like that's the general way I would characterize the the role that it plays I think there's a frontier of how these things intersect together like if you want to have a conversational modality on top of like your existing user data um which I I know for a fact like you know colleagues at various places that are doing this at scale are thinking about right now um I think there's some really interesting things we're going to see in the years to come um but we're we're very very early amazing well I'm very excited to see movement in in this space in the future and I'm excited we should probably have another conversation like this in in six months or 12 months and see it would be a very different conversation yeah absolutely things have been going yeah yeah um maybe we can even use gb9 to create the conversation um for us I am to um a let me so Myan did have a a really interesting question which I think um I think it's provocative um in a way that'll become apparent but um so Myan has asked considering we've already mentioned that AI hasn't reached a level of humor do we think um they would be able to film directors in the film industry particularly for comedy scenes so I I just want to preface this by saying um how do I want to say this maybe not directors I I don't know about directors I I don't know about writers either but I would argue that a lot of Hollywood scripts seem like they've been generated algorithmically for decades pre yes I've read some pretty horrible l in my time yeah no no for for sure I know it's it's a great point I think like it's so you know I think there's there's two answers people tend to have this one is like it's not going to replace anyone it's just going to augment everything and I think that's a little bit like of a of a false statement because yes it's it's there are certain tasks which it's absolutely going to replace and there are likely Focus f for folks who focus on those tasks today who will likely have to Pivot to something else example I give outside like I do what I do often say that um technology replaces tasks not jobs and it is our job as a society to figure out job tasks but having said that think about truck driving right which is not just a task but a job and that is something that it very much seems will be automated right once we have are well aware of this yeah that risk which is a source of huge employment in the United States in particular the economic impact of that is going to be yeah there's there's a lot of concerns there I can totally see it yeah I mean so so I think it's it's it's um it's flaw to pretend like it's not a a concern or a problem but I think like to your point like I think what we generally find in the long term is that you know these things play out positively like you end up with and I mean again you know mileage may vary but like generally you find like you know extrapolate T equals to Infinity people develop skills on how to use these new models in more effective ways and and and you know have even more productivity and better quality output like most disciplines of Art in my opinion have advanced positively the same time we have a lot of garbage I feel like the art today the highest end of it is better than anything that's been produced before um in pretty much almost in any field I would I would wager and I'm sure I'm pissing off some people by saying that um but but it's it's it is true to some degree um so I think like there's a that the challenge is sort of that change management and the piece like in between which is where I think you you really have the risk of of of turning things upside down for a lot of of folks um but I guess just practically speaking around like the film industry um my sense is you're you're likely already seeing people use it heavily as like an augmentive tool and an assistive tool um I think what you'll see is probably like greater leverage usually the way this goes is two ways one is it starts increasing leverage of of individuals which means they're either able to create stuff of Greater scale and greater quality than they were before or you're able to hire fewer of them which means job loss the other direction in cases of like where there's good employment protections as we found in the in the film industry um is that you often see entrance from other spaces start to take advantage of this so imagine ugc where there are no such concerns around jobs for the most part because they're these aren't traditionally seen as Jobs these are Hobbies right and and there's the possibility of like cannibalization if the quality of that stuff is able to get sufficiently high and and and you know actually start competing with premium video which it frankly already is in in many contexts in many days um so I think those are the probably the two Avenues you're going to see of impact but early days we'll see how it plays out yeah and to your point of there being and I'm going to paraphrase you did not say this but if to paraphrase and slight project there's a lot of crap out there right is my paraphrase and projection um having said that like printing press arguably one of the greatest Technologies ever created just imagine the absolute amount of nonsense that that te and the same is true for the the same is true for the internet like you think about Google and and I'm sure many of us have realized how much worse Google results have gotten in the last like six months to a year like with the kind of SEO optimized garbage that shows up at the top and it's not that Google has gotten worse it's that the adversaries have gotten better right exactly and this is and yeah it's a fundamental problem with any of these but yeah it's an ongoing yeah yeah so I'm really interested in just stepping back a bit I mean you know maybe a bit philosophically sociologically I like I do think there's a a serious false dichotomy between um creative fields and technological Fields I do think there are feedback loops which have meant that that has actually manifested deeply in society um so stem versus created um and what you're what we're told as children we can do and what should we should do and these these types of things um but I my provocative statement is that um the creative Fields perhaps the humanities more generally are in some sort of tech induced existential crisis currently um and and I'm interested in your thoughts on I don't I don't entirely believe that I think it's a c to no no it's a it's it's a fair thought and honestly I would I would absolutely agree that they they're in a crisis like I think like if you if you talk to anyone who speak who who lives in these spaces it's it's it's very present and people have very different reactions to it you know some are excited by the possibilities of of you know enhancing their creative abilities and enhancing their creative outputs and others are threatened uh justifiably so I would argue in many cases um the way I would view it is like I think again short-term versus long term long term I think there's really there's really two ways you can interpret this right two perspectives you can have one is you know this is the singularity right and and you know all human jobs all human creativity is going to be consumed by this thing sooner than we're able to adapt to it and you know that has massive consequences for society that go well beyond like the Arts and and and in that case I think you know all hope is kind of lost in some ways and we're we're yeah they the other perspective which I tend to hold more to is this is just yet another step change and we've seen step changes before and we've adopted to them before and we've become more productive and you know I think we've ended up in a better place I would argue and I would I would hope that we that this the similar story plays out here um the one which concerns me is really like the short term and what happens to folks in the next you know like five 10 years yeah because in in the short term we have seen at the same time as having um screen access Guild strikes writer strikes of course and this is provocative I don't necessarily want to focus on this particular thing because it's something the media generated as well but at the same time as that was happening we saw like a $900,000 job listing for an a product manager which which to be fair I I've heard from my insiders that had nothing to do with Gen it was like all like actual recommendations oriented stuff it was but it is it is a good it is a good valid concern and and I think the the the inequality piece of it is totally right on um I think like one way of also framing this is that you know I think like these changes happen these changes are not like gradually they don't happen gradually like in sense like when this new technology shows up it shows up and it changes overnight and humans aren't able to adapt on that time scale like we aren't able to just suddenly switch careers and build new skill sets and all of that um and so I think like ultimately like this is a a trade-off between speed of innovation and change and and buying people time to adapt U which you know in my mind that is actually what a lot of the I mean I can I'm I can't I'm speaking well out of my expertise but I feel like a lot of the the protectionist sentiments are really about just buying time for people to be able to adapt and and participate in this much like brighter future um so I don't see those two things at odds at all um it's just an uncomfortable kind of like battle that that you know is playing out yeah I agree and I do think it is a there's a certain jostle that happens in a period of transformation which can be uncomfortable but it does come down to what what we value as a society and where we what our value and where we put that value as well and I do think um we definitively um value the output of all types of creativity and it's will be very interesting to see how these things are are combined um to create works that connect us as humans connect us via storytelling um via sense of of humor um and allow us to utilize computation in in order to do that in really fun and exciting ways yeah um appreciate your your time fun so much fun thank you so much to everyone who joined and I tried to get to as many questions as possible but there were so many more so please do I've just um put the link to our uh slack uh workspace um in the chat again go to AMA guests and Ven and I will answer any any questions you you have um really looking forward to seeing those questions for what it's right yeah absolutely um yeah so please do do that and I'll go through and try to copy and paste a few of them as well but if you ask one that we we didn't get to um please please do that um Ben do you have any any final thoughts on um anything with this how people get started working with this type of stuff today yeah yeah yeah I mean I think one of the one of the Beauties is like in in ml and in technical spaces you often see these periods of kind of almost like a reset around skill sets where you know this like when when when deep learning f show up it felt like a reset for many folks in the ml community and the ban community because it felt like a lot of the skills they had built up were now suddenly no longer applicable and they're replac they were sort of starting from the beginning like everybody else and I think we're kind of going through another s such transition Point um and that's awesome by the way because I think it really opens up the floor to like a lot more people to engage from a lot of different backgrounds and I I I still think like some of the foundational skill sets are valuable here like around like like I was mentioning evaluation and and how to like you know procure data and think about it and all of that um but I do think like this is possibly one of the more accessible times in ml history um and and and yeah looking forward to seeing more democratization yeah absolutely and I do think there were several steps in the the Deep learning story right so firstly you had kind of the creation of deep learning Frameworks that abstracted away a lot of the linear algebra and and calculus yeah and back propagation so you could do that stuff without being a deep math Insider um but you still needed like some hardcore software skills then we had I mean Caris is one of the the first apis which I just thought did one of the most beautiful things to allow lots of people to approach it then of course pie torch took off we had The Descent of tensor flow and then pie torch of course requires significant boiler plate and significant in the end software stuff then you have Frameworks such as P torch lightning which allowed you then to work with pytorch in in in a different ways so it's a fascinating interactive space of all different tools Technologies and methods which in the end will make make these types of um Technologies accessible to to everyone and and for what it's worth I think we're still seeing like an active battle to determine what's going to be the winning Frameworks for this new kind of generation of models and use cases and and just having tried to deploy stuff in space this is not at all a settled question I think there's a lot to learn in the next you know coming days absolutely well I'm excited to have this conversation again in six or 12 months and see see where we're at yeah awesome well thanks once again Ven and thank you everyone for joining and look forward to chatting on everyone

Original Description

Vinith Misra is a machine learning scientist and software engineer at Roblox, an online platform for gaming, co-experience, and game creation that brings people together through play. Previously, he led the Artwork and Video Data Science team at Netflix, where they used multimodal (video/text/audio) machine learning to assist creators, and analytics to inform creative decisions. He has also worked as Research Staff at IBM Watson, after graduating from Stanford (PhD) and MIT (M. Eng). In a parallel world, he worked as a technical consultant for HBO on their show Silicon Valley, including “developing” the fictitious middle-out compression algorithm! ​In this fireside chat, Vinith joins Hugo Bowne-Anderson, Outerbounds’ Head of Developer Relations, to talk about the intersection of AI, machine learning data, algorithms, human creativity, content, and strategy in the entertainment industry. They’ll discuss - ​industrial applications of machine learning, computer vision, and AI to creative workflows and what happens when machine learning engineers work with content creators, such as directors, editors, 3d artists, and game developers; - ​Multimodal GenAI systems, where they came from and where they’re going, with a view towards both linear content creation and experiential systems, such as gaming; - Under-reported technical challenges in the GenAI space, such as developing robust evaluation systems; - Moving from ML and GenAI POCs to production and what is actually working in industry. ​And much more! 00:00 Prelude 03:03 The fireside chat begins 05:24 Introducing Vinith, machine learning scientist and software engineer at Roblox 07:54 The importance of humour in computation 13:32 Working as a technical consultant on HBO's Silicon Valley 18:40: What happens at Netflix when machine learning engineers work with content creators 22:48 Tools and techniques for ML meets creatives 27:41 How GenAI can change the game here 31:51 Multimodal GenAI systems for gaming at Roblox
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Playlist UU5h8Ji6Lm1RyAZopnCpDq7Q · Outerbounds · 50 of 60

1 Metaflow GUI for monitoring machine learning workflows
Metaflow GUI for monitoring machine learning workflows
Outerbounds
2 Metaflow Cards [no sound]
Metaflow Cards [no sound]
Outerbounds
3 Fireside chat #1: How to Produce Sustainable Business Value with Machine Learning
Fireside chat #1: How to Produce Sustainable Business Value with Machine Learning
Outerbounds
4 Fireside chat #2: MadeWithML.com -- Teaching Practical Machine Learning
Fireside chat #2: MadeWithML.com -- Teaching Practical Machine Learning
Outerbounds
5 Metaflow on Kubernetes and Argo Workflows [no sound]
Metaflow on Kubernetes and Argo Workflows [no sound]
Outerbounds
6 Fireside chat #3: Reasonable Scale Machine Learning -- You're not Google and it's totally OK
Fireside chat #3: Reasonable Scale Machine Learning -- You're not Google and it's totally OK
Outerbounds
7 Metaflow Tags: Programmatic Tagging
Metaflow Tags: Programmatic Tagging
Outerbounds
8 Metaflow Tags: Basic Tagging
Metaflow Tags: Basic Tagging
Outerbounds
9 Metaflow Tags: Tags in CI/CD
Metaflow Tags: Tags in CI/CD
Outerbounds
10 Metaflow Tags: Tags and Namespaces
Metaflow Tags: Tags and Namespaces
Outerbounds
11 Metaflow Tags: Tags and Continuous Training
Metaflow Tags: Tags and Continuous Training
Outerbounds
12 Fireside chat #4: Machine Learning and User Experience -- Building ML Products for People
Fireside chat #4: Machine Learning and User Experience -- Building ML Products for People
Outerbounds
13 Fireside Chat #5: Machine Learning + Infrastructure for Humans
Fireside Chat #5: Machine Learning + Infrastructure for Humans
Outerbounds
14 Metaflow Sandbox Demo: Free Data Science Infrastructure In the Browser
Metaflow Sandbox Demo: Free Data Science Infrastructure In the Browser
Outerbounds
15 Metaflow on Azure
Metaflow on Azure
Outerbounds
16 Fireside Chat #6: Operationalizing ML -- Patterns and Pain Points from MLOps Practitioners
Fireside Chat #6: Operationalizing ML -- Patterns and Pain Points from MLOps Practitioners
Outerbounds
17 ML engineering vs traditional software engineering: similarities and differences
ML engineering vs traditional software engineering: similarities and differences
Outerbounds
18 Why data scientists love and hate notebooks: velocity and validation
Why data scientists love and hate notebooks: velocity and validation
Outerbounds
19 What even is a 10x ML engineer?
What even is a 10x ML engineer?
Outerbounds
20 The 4 main tasks in the production ML lifecycle
The 4 main tasks in the production ML lifecycle
Outerbounds
21 Is the premise of data-centric AI flawed?
Is the premise of data-centric AI flawed?
Outerbounds
22 The 3 factors that Determine the success of ML projects
The 3 factors that Determine the success of ML projects
Outerbounds
23 Fireside Chat #7: How to Build an Enterprise Machine Learning Platform from Scratch
Fireside Chat #7: How to Build an Enterprise Machine Learning Platform from Scratch
Outerbounds
24 Run Metaflow on any cloud: Google Cloud, Azure, or AWS [no sound]
Run Metaflow on any cloud: Google Cloud, Azure, or AWS [no sound]
Outerbounds
25 Metaflow on GCP
Metaflow on GCP
Outerbounds
26 Fireside Chat #8: Navigating the Full Stack of Machine Learning
Fireside Chat #8: Navigating the Full Stack of Machine Learning
Outerbounds
27 How to Build a Full-Stack Recommender System
How to Build a Full-Stack Recommender System
Outerbounds
28 Modernize your Airflow deployments with Metaflow - zero-cost migration [no sound]
Modernize your Airflow deployments with Metaflow - zero-cost migration [no sound]
Outerbounds
29 Easy Airflow DAGs for ML and data science with Metaflow [no sound]
Easy Airflow DAGs for ML and data science with Metaflow [no sound]
Outerbounds
30 Fireside chat #9:  Language Processing: From Prototype to Production
Fireside chat #9: Language Processing: From Prototype to Production
Outerbounds
31 How to build end-to-end recommender systems at reasonable scale
How to build end-to-end recommender systems at reasonable scale
Outerbounds
32 Full-Stack Machine Learning with Metaflow on CoRise
Full-Stack Machine Learning with Metaflow on CoRise
Outerbounds
33 Natural Language Processing meets MLOps
Natural Language Processing meets MLOps
Outerbounds
34 Fireside Chat #10: Large Language Models: Beyond Proofs of Concept
Fireside Chat #10: Large Language Models: Beyond Proofs of Concept
Outerbounds
35 What even are Large Language Models?
What even are Large Language Models?
Outerbounds
36 How to get started with LLMs today
How to get started with LLMs today
Outerbounds
37 LLMs in production
LLMs in production
Outerbounds
38 Accessing secrets securely in Metaflow [no audio]
Accessing secrets securely in Metaflow [no audio]
Outerbounds
39 Fireside Chat #11: The Open-Source Modern Data Stack
Fireside Chat #11: The Open-Source Modern Data Stack
Outerbounds
40 Fireside chat #12: Kubernetes for Data Scientists
Fireside chat #12: Kubernetes for Data Scientists
Outerbounds
41 Behind the Screen: How Amazon Prime Video ships RecSys models 4x faster
Behind the Screen: How Amazon Prime Video ships RecSys models 4x faster
Outerbounds
42 Fireside chat #13: Supply Chain Security in Machine Learning
Fireside chat #13: Supply Chain Security in Machine Learning
Outerbounds
43 Quick Delivery, Quicker ML: DeliveryHero's Metaflow Story
Quick Delivery, Quicker ML: DeliveryHero's Metaflow Story
Outerbounds
44 Crafting General Intelligence: LLM Fine-tuning with Metaflow at Adept.ai
Crafting General Intelligence: LLM Fine-tuning with Metaflow at Adept.ai
Outerbounds
45 Fuelling Decisions: How DTN Powers Gas Pricing and Data Science Collaboration
Fuelling Decisions: How DTN Powers Gas Pricing and Data Science Collaboration
Outerbounds
46 From Kitchen to Doorstep: Optimizing Data Science Velocity at Deliveroo
From Kitchen to Doorstep: Optimizing Data Science Velocity at Deliveroo
Outerbounds
47 Building a GenAI Ready ML Platform with Metaflow at Autodesk
Building a GenAI Ready ML Platform with Metaflow at Autodesk
Outerbounds
48 Media Transcoding for 10 Million users and beyond with Metaflow at Epignosis
Media Transcoding for 10 Million users and beyond with Metaflow at Epignosis
Outerbounds
49 Telematics with Metaflow: How Nirvana Insurance built a large-scale Risk Estimation platform
Telematics with Metaflow: How Nirvana Insurance built a large-scale Risk Estimation platform
Outerbounds
Fireside chat #14: Generative AI and Machine Learning for Film, TV, and Gaming
Fireside chat #14: Generative AI and Machine Learning for Film, TV, and Gaming
Outerbounds
51 The Past, Present, and Future of Generative AI
The Past, Present, and Future of Generative AI
Outerbounds
52 Building Production Systems with Generative AI, Machine Learning, and Data
Building Production Systems with Generative AI, Machine Learning, and Data
Outerbounds
53 A Custom Fine-Tuned LLM in Action (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 5)
A Custom Fine-Tuned LLM in Action (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 5)
Outerbounds
54 Building Live Production Systems with RAG (LLMs & RAG: An Interactive Guided Tour Part 4)
Building Live Production Systems with RAG (LLMs & RAG: An Interactive Guided Tour Part 4)
Outerbounds
55 Better Relevancy with RAG (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 3)
Better Relevancy with RAG (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 3)
Outerbounds
56 Working with OSS LLMs (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 2)
Working with OSS LLMs (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 2)
Outerbounds
57 Hitting OpenAI and Other Vendor APIs (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 1)
Hitting OpenAI and Other Vendor APIs (LLMs, RAG, and Fine-Tuning: An Interactive Guided Tour Part 1)
Outerbounds
58 Production Systems with Generative AI (LLMs, RAG, & Fine-Tuning: An Interactive Guided Tour Part 0)
Production Systems with Generative AI (LLMs, RAG, & Fine-Tuning: An Interactive Guided Tour Part 0)
Outerbounds
59 LLMs in Practice: A Guide to Recent Trends and Techniques
LLMs in Practice: A Guide to Recent Trends and Techniques
Outerbounds
60 Metaflow for distributed high-performance computing and large-scale AI training
Metaflow for distributed high-performance computing and large-scale AI training
Outerbounds

This video explores the applications of AI, machine learning, and computer vision in the entertainment industry, with a focus on generative AI and multimodal systems. It discusses the intersection of technology and human creativity in content creation.

Key Takeaways
  1. Understand the basics of computer vision and machine learning
  2. Explore the applications of AI in the entertainment industry
  3. Develop multimodal systems for film, TV, and gaming
  4. Integrate AI with human creativity in content creation
  5. Build machine learning pipelines for creative workflows
  6. Evaluate and refine generative AI models
💡 The key to successful AI applications in the entertainment industry is the integration of technology with human creativity, and the development of robust evaluation systems for generative AI models.

Related Reads

Chapters (8)

Prelude
3:03 The fireside chat begins
5:24 Introducing Vinith, machine learning scientist and software engineer at Roblox
7:54 The importance of humour in computation
13:32 Working as a technical consultant on HBO's Silicon Valley
22:48 Tools and techniques for ML meets creatives
27:41 How GenAI can change the game here
31:51 Multimodal GenAI systems for gaming at Roblox
Up next
9-Phase Computer Vision Roadmap 2026 | AI & Deep Learning | #shorts
SCALER
Watch →