Building Multimodal AI Apps with Replicate — Full Session

Replicate · Intermediate ·🧠 Large Language Models ·1y ago

Key Takeaways

This video teaches how to build multimodal AI apps with Replicate, including a movie generator app

Full Transcript

So I put together a few slides just to kind of demonstrate all right some of the capabilities of the platform um which you already touched on Hugo um but there's there's so many things you can do with replicate but I kind of wanted to boil it down into four main things. So first obviously is run AI models. You don't have to um you can run AI models on our website itself which is pretty revolutionary like you don't have to go and write code. You can go straight to the platform and run a model. So actually let me just show you one of the models. This is an example model page uh you would see on replicate. And it's kind of like we structured it like GitHub where you have all these repos um but instead we have these different models or functions that you can run. This is our Nvidia Sauna Sprint um model that we have here. This is actually I believe this is our fastest image model on the platform. Um but I just want to show you how easy it is to run something. Um so I'm just going to type in a very arbitrary prompt like a person eating grass. Let's just give it something weird. Would you mind just um maximizing a bit so we can zoom in or Yeah. Can you see that? Yeah, that's Wow. Yeah. So, that was super fast. It's a 0.1 second runtime for this model. Um but literally, you just get a stunning image in in less than 0.1 seconds on our platform. You can just run it straight. Um, we could do like an elephant walking through the mall, for instance. That's another prompt. Hit run. Easy. Super easy to run any ML models. We have a bunch of these kind of similar models on our platform that you can go and explore. Uh, Hugo, you demonstrated the explore page. You can kind of just go here, click image models for instance, and find like a variety of things. We've tried to really curate stuff to find something for your specific use use case. Um, so that's quickly index on the latency that how quick that was because like for anyone who's tried to generate images in chat GBT recently or Grock or whatever it may be, like amazing capabilities, but it it takes a while to be honest. Right. Totally. Um, you know, 4 O's image or ChachiPT's image generation is incredibly powerful. It's auto reggressive. It's an amazing technology, but um, it's it's difficult to integrate it in apps where users want to have fast latency. And so something like this, this particular NVIDIA sauna sprint model is huge for people. It's already generating crazy good images. So this is something you can actually go and use our API for and integrate it into your own projects. Um so that's running AI models. Another thing you can do is fine-tune models. So um I want to get a sense of how many people are familiar with the process of fine-tuning or like Laura's um is that something that like a term that strikes strikes something for people? I'll just ask in the discord and depending on appetite next week in in in the workshops where we'll be doing things that people really want to focus on. So we may be um going deeper into those things. Um, I'm going to ask how many people know about Laura essentially. Yeah, we have have several more. No, that it looks about 50/50. Take into account. Okay. Okay. Pretty cool. Pretty cool. Yeah. So essentially for for those who don't know fine-tuning um models just basically means that you're able to provide an already existing base model some extra training data and you can actually create forks of models or variations of models that will give you pretty specific outputs. So for instance with image models you could give it photos of blueprints for instance and then create your own image model that will give you outputs that look like blueprints specifically and we have functionalities on our platform. So, this I'll be diving a little bit deeper later, but we have a Laura trainer where you can just give it a bunch of images and you'll be able to create your own version of an image model that is specific to um your training data that you provide. And again, you don't have to know the machine learning uh intricacies behind this. It's literally just our simple web UI. You can provide it your input images and then hit run and then you'll have your own fine-tuned model. So it's very very simple, very straightforward. Cool. So that is fine-tuning models. Another another thing you could do and and something that you guys are probably pretty interested in is using models in your own projects. So every single model on our platform has an API and you can actually call this API within your own code, your own projects and this is how people are able to make really really cool amazing projects with replicate. um just from a simple API interface, we've been seeing some crazy crazy cool things that people are building both in the software realm and in the hardware hardware physical realm with Replicate. So it's just really cool. We've built this amazing community of indie hackers and builders um using our platform. And then finally, you can publish models. So if you actually are savvy with ML engineering, if you know how to create your own models or even um the again the misconception is you don't even have to run code that involves a GPU. You don't have to make generative image models or generative text models or generative video models. You can run functions. You can run code and actually call that as an API. So let me demonstrate. We actually even have this as a model. We've classified this particular function as a model. And it's literally a hello world function. All you have to do is type in a name. So I'll type in my own name, shreder. And then when you hit boot and run, this will actually just provide you um and again this is not GPU based. It's just a CPU function. You get the actual function. It's just a simple Python function um that you can use. Even the goo, we call this like background that's part of our brand name goo. Um this is also a function that we have on replicate that you can use and you can create MP4 videos with it. You can create PGs, but this is not something that requires a GPU. This is not an AI model. This is a straight up CPUbased regular plain old code function. uh like you can go to the GitHub right here and you can see all the code associated with it and it doesn't require a GPU. It's not a generative AI model. Um but it's an API on our platform because you can use it as sort of like a serverless API function and integrate it into your apps. Right? So those are the four things I've boiled it down to with what you can do with replicate. Um again running all the models we had our model curation um all these different use cases but something else we have is all of these model providers like claude anthropic we have comfy UI we have flux which is one of the leading image generation models we have stable stable diffusion from stability AI we have meta's language uh models which are the llama models we have some really cool video models so we have cling which is which is going really popular right now. We have WAN, we have uh Luma, we're getting Pika on there, and we even have DeepSeek. So, all of these models that you know, they're all on replicate, and you can use them and call them in your own in your own projects using our API. Again, you can fine-tune models. It's an easy web UI, no code. Um, we have image and video fine-tuning. So all you have to do is just go on to our platform where we have a fluxdev Laura training. Um and again I'll be sending like all the all these links to Hugo so that he'll be able to disperse everything to you guys. Um but again this process of creating your own fine tunes is very very simple. You can just curate a set of images and pass it into our our trainer and then you'll be able to create amazing fine-tunes. And let me just show you specifically for video fine-tuning what I kind of mean. I know it's kind of difficult to understand like what does it mean like if you provide a bunch of images to an already existing model like what can you make? Let me show you some of the uh fine tunes that I made with WAN. So WAN is a uh video model that has recently come out um WAN 2.1 um and this is our trainer for it. So essentially what you want to do is you'll provide it a destination. So I already have like all of these different WAN fine-tunes and then I provide it a list of images that I kind of want the video to mimic the style of. So I've I've tried like you can see blueprints, I've tried cinematic, I've tried Studio Ghibli which is a huge trend. um and like all these different types of styles. Um and you'll just follow all the steps here and you can create your own video model. So I've created like a Studio Ghibli version. Um so this is again the WAN uh video model but fine-tuned using Studio Ghibli images. So you you're able to create videos like these. Um and you can see like some other examples here. Like we have I kind of did like a life of pie one. Um, we have a tiger and like a in a boat in Studio Ghibli style. Um, and then like I don't know this one this one is is one of my favorites. Just a giraffe with a young boy eating some burgers. But you know this it's it's super simple. You just pass in a bunch of images and then you're able to train it. And again, you don't have to know the machine learning that goes behind it. It's very complicated, but we take we take care of all that for you. So it's very cool. Um, and I have another one that is a Van Go style. So, you can see I've created some like Van Go videos with WAN. All I did was pass in some examples of Van Go's paintings into this model and fine-tuned it. And you can see I've been able to create some really, really cool examples um, just with simple text prompts. Um, I'm getting Van Go style videos. Um, this one is one of my personal favorites. a giraffe walking through the sundrenched streets of San Francisco. And it's it's pretty good. I really really like the style that's that's been able to come out of this. Um, you know, and you can create create very similar styles. So video style transfer is now a thing with replicate um and video fine-tuning. So that you can easily do on our platform as well. Um, all right. And fine-tunes our model. So you can actually go and use one of my own models that I've created or create your own in your own projects. So it has an actual API associated with it. The for instance the Jibli video model that I created has an API that you can go and use in your own projects. You can call it in you know a web app, you can call it in a mobile app or anything. Um so fine-tunes are just as equally models as you know something that you put um you write code for or um any of the other models on our platform. And these are some of the examples of image fine-tunes uh that we have curated on our platform. So you can see like some people have made Mona Lisa style. So you can uh type in any prompt and you'll always have Mona Lisa's face in the image. Uh we have some really cool like black light is one of my favorites. Um so you create like a black light kind of neon style image. Um we have like Spongebob, we have uh cinematic style. Um so these are all just like example um flux fine tunes uh that people have created. And again you can actually go ahead and use these models in your own projects. And um it's really cool. We've we've gotten some really creative um just kind of flux fine-tuned style stylistic image models uh just from people like uploading a bunch of images and then creating these models uh very easily very um from scratch. Um so it's very very cool. And then the thing that most people are really excited about with Replicate and what you're probably excited about is being able to make things like what exactly can you make with all these APIs that are on our platform. We've had some amazingly creative projects come out of replicate um or indie hackers building on top of replicate. So you can see um the CEO of Verscell uh Gummo created a Airbnb clone using his uh Vzero product. Um so Vzero is like a code editing tool and he asked to use replicates fast flux which is an image generator generating model to provide images for all these Airbnb listings. So that was pretty cool. Um we had this goo configurator. So again, like the goo that you see in the background of everything that we do is also a web app. Um, someone made a command line interface with an integrated LLM. So you can actually just like type in what is the command for like searching for this file into your terminal. Um, and it's LLM powered will give you the command all powered with replicate. Uh, other people have been able to make and scale um, businesses with replicate. So we have um restore photos which you see down there. Uh that uses one of our restore photos APIs. Um you also see um the interior design AI app that uses one of our models that was created by levels io on Twitter. Um so he was originally using uh one of our uh interior design APIs and he's actually been able to scale that into a a pretty substantial business. Um so it's very very cool. We also see like physical AI. Someone has created a robot with uh some of the language models that we have on our platform. So, llama, lava, whisper for speech to text. These are all plat these are all models that are on our platform that you can just like link and chain together and create some really really cool things. Uh, so we've been able to see some amazing creative projects come out of the community that we've built on Replicate and it's just absolutely incredible and we just kind of want to inspire more people to use the platform and create some of these insane projects that we never would have thought of just with the thousands and thousands of models that are on our on our platform. Um so yeah I actually kind of wanted to walk through a example of something that you could create with replicate. Uh so we have again uh APIs associated with all these models. Uh you can use APIs with Node.js. So React apps you can use it with Python. We have a Python client uh HTTP. So you can do like simple curl requests as well to test out models. So these are like the three main ways you can integrate APIs into your uh projects. So I wanted to walk through a example repo that I literally created in full discretion. Literally created two hours ago. Um and it's it's insane what you can do with replicate and AI code editing. So again, full discretion. Um this this is a movie generator app that I created with Vzero. Um V 0ero is a code editing tool. Um and I created this literally two hours ago. So um yes. So I'm going to show you exactly like how my my entire process. I have my entire uh Vzero hit uh chat history. So we can sort of expose um what I was zooming in a bit and maybe removing that turn on reactions just to the so that we can Yeah. Can we see now? Yes. Yeah. But if you could full screen it as well this this window, right? Yeah. Yeah, that would be great. All right. Um how about we like fully full screen it? Yeah, that'd be awesome. And feel free to close the lefth hand panel if you don't need it as as well. Yeah. How's this? That's awesome. You could even close the chat thing on the left as well. Sorry, where it says new chat on the left. The chat. On the left up the top in the top. On the left. Oh, okay. Yeah, I didn't. Okay, cool. All right. Well, anyone who hasn't played around with Vzero, definitely do and I'll link link to it. They've got a lot of cool stuff. Yeah, Vzero is is incredible. So is cursor. So is lovable. All of these AI code editing tools I highly encourage everyone to play around with especially because they have knowledge of replicate within the LLM. So you can actually just like say use this API and it will actually create something with this API. Um so I'm going to expose like a little bit of my my thought process here. So what I'm I'm creating here is a movie generator app. So I want the user to input a prompt of what exactly I want to see. uh in my little short film and then I also want to provide it a movie to stylize it from. So like I want something for for instance uh I want to create something in the style of Indiana Jones or I want to create something in the style of Inside Out or I want to create something in the style of Studio Ghibli. Um that is what we're going to do. So, uh, the first thing I just started off with, um, I just said create a simple modern dark nex.js web, uh, UI for this. Um, and we're going to like list our inputs. So, it's going to be the video prompt and the inspiration film. So, um, you can see like this is what we have for our entire web app. So, we have our video that's going to show up here as well as our two prompts. So, we have video prompt and inspiration film. Um, so you could see I was like kind of yelling at it a little bit. There was like some moments where it was it was frustrating me. So I wanted to get rid of the title. Um, but you can literally just tell it exactly what to do. So the the most creative but I would say like the most like lift part of all this is that you have to actually find the models that you want to use. So in this case, what I'm doing is taking text to image to video. So I'm going to take the the user's text prompt that they provide in video prompt and inspiration film and create a starter image for the video to start from. That's going to be like the first frame. And then I'm going to pass that image that I create to a video model that will generate video for me. So, one of the best image models I found, especially for anything that requires like like maybe IP or like you know, we're dealing with movies here. So, I've actually been able to find that Google's image and three model which is on our platform is able to handle things that like are movies or like books or like anything that can uh really represent the style of something. So like movies um we can see we can go up to I'll show you that's just a preview of what you'll be seeing but we have Google image and 3. Um this is a beautiful model. It's very fast. It's very capable. Uh you can see like this is the sample prompt that we showcase um on the model page for Google image 3. Um and it's able to create a strawberry hummingbird which is very cool. Um, but this is the model that I'm going to be using to transform the user's um, video prompt into a starter image. Um, so we can take a look at let's see, we can take a look at let's go back to vzero right here. So what I'm asking it to do is first of all create the API endpoints that I need uh, to be able to create this this app. Um, so I'll show you in the code here. We have two API endpoints. Um, and I'll I'll just sort of structure this like maybe um, we we're intro to Nex.js or React. But essentially what we'll have is a front end which is our components. We have a movie generator component which is actually rendering what you see on the front end right here. Um, so this is what's showcasing the actual video that we're going to be seeing. Um, and then we have our API endpoints. So you can see in the actual code, let's let's see, can I zoom in? I'll just highlight. Um, we have step one, which is generate image. So what we're doing first is taking the user's video prompt and creating a image with this. So we're taking again remember we had a video prompt and an inspiration film for our two uh text prompts and we're passing that into our API endpoint which is generate image. So let's go ahead and go into generate image here. This is our API endpoint which uses the replicate um model that we have right here. And actually this should be version six which is the same thing. So generate image. What we're doing here is taking our Google image and 3 model. You can see right here that's what we're using. And we're also taking in our input which we have constructed right here. So you can see this is what I'm constructing. I'm prompt engineering a little bit my prompt that I want to provide to Google image 3. So I take the video prompt and then I say comma in the style of the film inspiration film. So for instance, it would be like I've been using a duck walking in the mall, in the style of the film Indiana Jones. That's sort of how you would construct the prompt that's you're passing into Google image 3. With replicate, you will, you know, it's a typical client. You're going to have an API token which you'll go on the platform. I can show you. We have API tokens. You'll just go in. You'll see some of your API tokens. You can create a token. you'll call it something. So I can say like test one create a token and then you can use that copy your token and that's how you'll interact with replicate in your in your code. Uh let's see we'll go back here. So this is what I'm doing right here. I'm creating the replicate client. Um and that's how we'll instantiate all of this code. So the next thing you want to do is actually run your prediction. So we'll take our input uh which is our our video prompt and also our model which is Google image 3 and then we create a prediction with that. Um after we do that we get our final prediction. This is what I'll return back to um the movie generator. And I'll be sharing all this code uh with Hugo so that you guys can see it as well as there's a bunch of really good information in our docs on how to get started um with uh you know setting up the API and setting up uh the the syntax of it all. It's very very simple. Um but just for the sake of time and showing you how all this works together, we're going to be first getting our image response. So this is again what's passed what's outputed from Google image3. So this will be like our starter frame. And now we want to pass this image into our video model. Um and this will be like the sort of starter frame for our video model. And so you can see right here we have our image URL that we obtained from Google image 3. And now we're going to generate our video. So we have a generate video API endpoint as well. Um, and again, same thing with the generate video. Uh, we're going to be using WAN. So, if you remember, I said, uh, WAN is a very, very capable video model. Um, not a lot of people are are really talking about it, but it is very very fast. It's pretty amazing what it can do. Um, the biggest thing is that it's very fast. I really like how fast it is. Uh so in particular I'm using WAN 2.1 image to video 480p and we can go on the replicate platform and I'll show you exactly what the model page looks like. So we have our WAN 2.1 I2V480p. You can see we pass in a uh input image and a simple prompt a woman is talking. And this video in particular was generated in under 40 seconds which is very very good for video generation. Um and you can also customize you know all these different parameters. So number of frames you can do frame rate frames per second. Um you can even add luras. So again like the fine tunes uh that you create. So maybe you did a studio jibly fine tune. You can add the weights here. And for anyone who's very very niche and specific about Lauras, you can even add uh Civot AI URLs or you can add hugging face URLs as Lauras into um these videos and create very stylized videos. Um so very cool, very uh tunable. You know, you can very get very very creative with this whole process. So let's go back to our code here. So again, same thing. We're creating a prediction with this particular model with this WAN model getting our prediction ID and then that will be our finalized output and this is what we display to the user. So if we go back to our our movie generator component which again is what is demonstrated in the front end um we will be showcasing that video URL in our in our front end. So, that is pretty much it. And again, like I coded this like two hours ago. Um, and it's it's working great. So, we can go and see I'll show you like an example. So, I've been doing a duck walking in a mall and inspiration film Inside Out, which I really like that film. So, um, I'm gonna hit generate movie and we can go to the replicate dashboard. So you can see that image and 3 is running and that is from again our API call and you can see instantly we got our first image which is our our key frame that is going to be sent to the WAN model to start creating our our actual film. Um so we go back to the dashboard. Now you can see WAN is running again from the API endpoint that we created and you can see all the logs. So again, the WAN model is going to create a video starting with this frame um and then output into our web app. And there you go. So we have a we've created a movie generator with two replicate API endpoints that we've chained together um for the user to simply just type in a video prompt and an inspiration film. Um and even even Vzero added download video sharing. So, I didn't even ask it to add that. So, again, highly encourage people to use these AI code editors because they're incredibly powerful and especially with replicate um models. Again, the the most difficult but the most fun part is actually discovering the models and chaining them together to make them create value that you never would have thought is possible. So, like we've created this um movie generator in the span of two two hours, a few minutes um and it's it's doing amazing. I want to see if maybe there's some someone in the audience who wants to give a video prompt and an inspiration film. I would welcome someone. I'm happy to, but I' I'd prefer someone else else to. Yeah. Yeah. Anyone in the Anyone in the audience? The first inspirational Yeah. Great. Go, Caleb. Uh, Eternal Sunshine of the Spotless Mind. Okay, that's a that's a good one. Eternal [Music] Sunshine of the Spotless Mind. And should we keep the same video prompt? Do we want to switch it up or All right, let's let's keep the same prompt and see what we can do. So again, let's go back to the replicate uh dashboard and actually see these APIs at work. Um so again, you have a very clean dashboard that shows you all of your runs of every single model. So we can hit the image in three. Okay, pretty interesting. um not very eternal sunshine specific. I would want it to be like more kind of like blue or um and it is a rubber duck. So, uh with more prompt engineering, again, this was just me saying blank in the style of blank. So, you can really get creative and experiment more with the because I do like Eternal Sunshine, but it's actually full of different styles as a movie itself, right? Yeah. Yeah, definitely. So, saving private Ryan. Oh, okay. Yeah, that is definitely very still a duck in a shopping mall. A duck in a shopping mall. All right. So, let's we'll have this run. So, you can see our WAN model is running and it's also in our UI showing that in the front end we're just loading everything up. Oh, that's pretty cool. Amazing. Oh, wow. And again, like there was no prompt engineering with the actual prompt. It's just a duck walking in the mall. Um, but WAN is able to create amazing video uh on our platform. And you know, again, it's just two API endpoints that we've chained together with replicate. Uh so let's try saving private Ryan generate movie and let's go to our dashboard and see if these APIs are running. So first again is the Google image and 3 model and you can see how it's constructed our prompt a duck walking in a mall in the style of the film saving private Ryan. We got some saunders in the background I guess. Um, but yeah, I mean I think it's pretty good. I think it's it's very fast model, very capable. Um, let's see. And to your point, there was no prompt engineering as like this is a like a this is amazing for a first pass. Yeah, totally. And again, I I coded or not even coded this, I prompt prompted Vzero like two hours ago to create this this demonstration. Uh, so it's already amazing what you can do in the f in the span of this is terrible walking. But again, with with better prompt engineering, this would just instantly get better. It's already amazing what it can do with very very simple prompts. Um, and this is not even the highest quality WAN model that we serve. This is our 480p model. We even have a WAN um 720p. So we have this curated collection of our videos with WAN 2.1. You can explore all of these. We have text to video. We we have text to video 480p, 720p, we have image to video 480, image to video 720, and we have WAN with Laura. So you can create stylized video as well. This particular one has the style of like flat color 2D animation. Um so you can really really get creative. Um, we have in painting with Wan. So, you can like paint over a certain uh subject and like get, you know, this is like the famous meme, but they've created a robot from him. Um, you know, so there's like infinite possibilities with not only the models you can create, but like stuff you can build. Um, we really want to see people um creating really amazing stuff. Replicate really values the community and like what we can do, anything we can do to support builders and get them really inspired to create amazing things. Um, but yeah, again, that was like in the span of just two hours of of prompting vzero with replicate APIs. Um, very simple to get started and yeah, this is pretty much it. Incredible. Thank you so much, Shuda. I um, we have lots of wonderful questions. Jas just asked, "How much would it cost to create those those three videos?" Sorry, I'm going to come back to Zoom so I can see everyone's face. Um, all right. Uh, how much would it cost to create those videos? That's a very good question. Uh, let me go back to our WAN models. So, Replicate is very transparent about the pricing that's on every single model. So, we can just click for instance this text to to video 720p. Um, this particular model is 24 cents per second of video. And I think this is automatically set at 5 seconds of video that you create. um which in the video world might be considered expensive but um again WAN isn't is probably the state-of-the-art video model right now but we have other video models that you can use that are much cheaper uh so we host uh Luma we host um I believe we have cling we have cling 1.6 six, which is also very capable and that includes uh camera controls. So, you can actually like tell it, you know, zoom in on this part or like zoom out or whatever. Um, but again, yeah, we are super transparent about pricing. We have the pricing up here as well as down here, you can see, you know, for 4 seconds of video, you're paying $1. Um, and we just want to make sure that people understand like the proportions and like what you're paying and all that kind of stuff. Um, yeah. Does that answer your question? Definitely. Super cool. And I I do want to clarify as well that like I can look at billing in my replicate account and it's super transparent and I can see what's up. And as someone who first encountered, you know, cloud billing costs on platforms like AWS, I'm eminently grateful for all the work you've done to be able to serve easily. And actually I I also think this is I I actually I think I'm not even intending this as a comparison because it doesn't do justice to replicate or hugging face. But I do want to say that um the hugging sorry the replicate dashboard and findability and explorability there um has solved for me an issue I have with platforms like hugging face which are incredible but it's not always easy to like it's it's chaotic right um or can can be chaotic let's say and my intention is not to cast shade on hugging face of course but replicate has made it beautiful to to explore all the different capabilities I also am interested we've There have been people chatting on Discord. They haven't used the term vibe coding. Um, which I'm a fan of not using that term actually, but um I So, we have been talking about vibe coding through this course and we we've got a session later this week on what we're calling software composing, not vibe coding, because I think that's probably a more useful um and less pointed term. Um, and there is a huge kind of space between writing every line of code yourself and not understanding every line of code yourself, right? And I always say that we've been vibe coding for years, just not at scale because yo, the amount of lines of code I pasted from Stack Overflow over the years, if you're telling me that I've understood them all, you know, I' I'd need to push push back on that. But I am interested um just in in your experience and how like you know the replicate API clearly but you'll use a tool like v0ero to supercharge yourself um and how how much of the code you kind of know works others you're debugging and then maybe you get frustrated and your sense of frustration when chatting with an AI assistant you are very gentle with it compared to some of the ruts I I I get in I can I can tell you that. Um but yeah, how how do you approach using AI assistance for for code generally? Yeah. Um that's a great question. I feel like after the release of cursor in August, uh for me everything has just been supercharged. Uh I I I mean I I went to school for computer science and I I was really instilled of like you know writing things from scratch or doing the Stack Overflow thing and just like really um sort of debugging from from what's already available. Uh but at some point, you know, I think it's not there's there's a sort of line in productivity. You know, are you really productive if you're trying to debug something that isn't really necessarily allowing you to understand your customers or understand the market? Those are those are really things that have been unleashed with the power of AI code editing tools. people now have the ability to really focus on things that matter which are the product, the UI, the customers, the use cases. Um, people have been able to experiment more with AI code editing. Um, so I say if you are, you know, kind of struggling initially when these tools were coming out, I was kind of like, you know, it's kind of weird that I spent four or five years of my schooling learning how to do all these things and now it's kind of just like gone away. But that's really not the case. um we've unlocked a new set of skills for people to play around with. It's, you know, a much more creative skill set. And so I really encourage everyone to use things like VZero or um you know, figure out the the niche prompts, figure out the things that really get extract the most value out of um you know, cursor or lovable or or cur um v0ero. And that will really allow you to just like focus on the things that really matter and the things that make you the most creative. I appreciate that so much. And Priya actually had a fantastic question, a specific question which I kind of want to zoom out from. Pria asked, "Anybody has anybody used lovable.dev?" Uh Pria wrote, "I have I built a full LLM app prototype with great difficulty. Not cool. Anybody have experiences to share?" So, thank you for sharing that experience, Priya. Um I've used Lovable a bit. I've used Replet a lot. I actually I was on the bus the other day going to the beach and used Replet on my cell phone to build a Space Invaders game that I played in real time on the way to the beach. Right. So that that was incredible. I've tried Bolt as well and all so the ones I'm mentioning lovable replet Bolt I I've used um for fun essentially. I've never tried to build anything serious with them. Having said that um cursor is an absolute game changer as you mentioned. So cursor in agent mode, not in yolo mode. Um, and we'll go into some of these things late later in the week, but using 3.7 Sonnet Max or Gemini 2.5 Max has been absolutely su supercharge me. But but I would only use it in I only use it really in two situations. One is if I know the technologies I'm using, like if I'm building LLM powered stuff and I know how to build those things, I'm okay doing it. Or if it's something I can debug in another way and I don't really care about the code that much. And the example I give there is building bespoke JSON viewers um for for other types of software I'm building. It does all this like React stuff that I don't really know about, but I don't care about either. And if I can run a few tests to make sure it's ingesting the JSON correctly and that type of stuff, that's all all I need there, right? Um but I do think you can get stuck. So if you know what you're doing or you don't really care what you're doing, they're the two cases I think you can be super superpowered with these. If you care what you're doing and you don't know what you're doing, um, you like these things. I need to pull them back. I know I'm babbling a bit now. I'm I'm I'm the problem that I'm talking about. Curs agents. They will write like sprawling code bases with subdirectories with thousands of lines of code and then you're going to have to be like, actually, can we move back and you give me the MVP in a 100 lines of code and this type of stuff as well. So, starting to learn how to work with these systems. Also, um, make sure to protect your API keys. Devon will go and use it to like really like Devon for me the other day like went and tried to do a 100 simulations with an API key when I told it explicitly not to, right? Um, so those are a few kind of rules of thumb, but I I wonder from your experience, Sha, um, what you're comfortable using these things for and what you wouldn't, if you have any heruristics for thinking about that. Yeah, that's a good question. And I mean, like I like I said, I feel like I use AI code editing like almost every single time. Anytime I'm doing code now, it's sort of second nature to me at this point. Um because again, it's been able to supercharge me and um really allow me to focus on things that are maybe a little bit more technical, like things I never thought I could uh learn or things I used to think were out of reach for me. Um so like building huge scalable systems or like um you know, anything like that. But I do agree that you know there is a point where you can't just like type in a prompt and like expect production level ready code. Um and that's sort of the thing that I feel like replicate is is really great at bridging the gap for for these kind of people. Um again we have you know thousand tens of thousands of models that you can explore and like really get creative with but we have an API associated with each and you know you can't go to cursor and be like build me a video generator a movie generator um you know it's not going to know the granular steps and so that's why I really encourage people to you know get familiar with all the models that are out there. There's so many text, video, image, uh, audio, 3D, even we have 3D models on our platform. Um, get familiar with all of those models and learn how you can sort of chain them together because cursor lovable replet, they don't know how to chain things together. The agent functionality hasn't gotten there yet. um you know so you that's sort of where you have to take the work and be like you know I I need to understand what what all the models there are out there and you know connect them with the associated input outputs um to get exactly what I want and get super super hyper specific totally um and I think getting to know the models in the systems is such good advice and that helps you discover antiatterns as well like I me I mentioned the scrolling codebase, right? Some of you may have encountered that when working with AI agents. Another one is called dead looping, right? And if you haven't encountered this, you will where there'll be a bug and you'll ask it to solve it and it will try to solve it and then there'll be another bug. Ask it to solve it. After several iterations, you get back to the first bug and it takes you on the same cycle and then after a few times you're like, "What are you doing, man?" And of course, that's when I start to use probably stronger language than than you used with Vzero, but you start to understand. Yeah, the the anti patterns also know that these things are improving on a daily basis as well, right? So what they look like now, it'll be very different in a few weeks potentially, if not, you know, a bit a bit longer. Yeah, totally. And I'll add that, you know, because the AI models are developing and these companies are really hyperfocused on making these models amazing. Like we've seen with WAN, video is is stunning already and it's just going to get better. the real value will be in builders like yourself um and everyone in the audience who are creating amazing UIs or amazing products around these APIs um because that is you know sort of the thing where you can just like if you're using a video model in your project you can just swap it with the next video model that's better than the previous one that you you created. So again I I highly encourage people to get really comfortable with building amazing products and building for users and understanding what users want. Um, I think that sort of is the skill is the skill shift that's going to happen. People are going to have to start understanding um, building for users and building incredible products. Totally. Um, there's one other thing that I think was implicit in everything you've demoed here which I just wanted to speak about briefly is composability. And I'm a hacker at heart, right? And Unix philosophy is one of the greatest things in the world as far as I'm concerned. the space of tools currently um and the ability to compose different tools and create chained workflows. And in our next workshop, everyone in in the course will be going through composing different things to move along the agentics continuum to more agentic like like um uh software more generally. But a platform like replicate allowing you to combine LLMs and um text to image, image to video, then including speech to text. So I could everything you just did we could just speak to our computer like using whisper or super whisper or using something in in replicate as well. So it's a fascinating space to kind of connect all different types of modular tools to build build really fun things. Yeah. And I'll say we have something very exciting in the works um along that theme. So keep your eyes peeled with replicate. Well that actually leads to my last question. I don't want you to give away any trade secrets or tell us how you know Google what Google spam filter is but is there anything that you you're excited about at the moment with respect to the future of the space that you can share? Yeah, I can say like the the all the models are just going to get way better. Um we're going to have more models and I think it's going to become a question of curating them and allowing people to really understand like you know which are the best for my particular use case. um it's getting insanely good and so again like I'm really really excited for the builders and the people who are able to extract value from these these models from the things that they've built. If I I again I highly encourage people to use AI code editing but also another tool is Twitter. Twitter is huge for understanding the space of what people are building and what these AI labs are creating. And it's insane to see what people have been able to make. Um, especially people who never coded before and are using AI code editing tools. These kind of these people are some of the most creative people and they've been able to create things that you wouldn't have thought are possible um with or without Replicate. It's it's absolutely incredible and I'm just really excited for the space and you know the whole democratization of coding. Um that's what I'm really looking forward to. Amazing. Well, thank you so much for sharing everything you're working on at Replicate and all the wonders of the platform. I'm excited to see everyone build with Replicate. Now, I think the only challenge is is is is time and being able to um find the time to try all these new incredible uh tools, but excited to see what you all build. Um and thank you once again for your sponsorship and and your time. Really appreciate it. Yeah, of course. So, everyone gets $100 worth of credits. Um this is much much more than enough to get started with the platform. Uh highly encourage everyone to take uh great use of it because there's so much to explore. There's so much to test and I want to see what everyone builds. So, amazing. Well, thank you once again and excited to see what you all build. Awesome. All right, see you everyone. Bye everyone.

Original Description

From the Build with LLMs course hosted by Hugo Bowne-Anderson on Maven. Shridhar shows what you can do on the Replicate platform and builds a cool movie generator app in less than 10 mins.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
Integrating Open-Weight LLM APIs: A Developer's Guide to Accessible AI
Learn to integrate open-weight LLM APIs into your applications for accessible AI, enabling you to leverage large language models like Llama and Mistral
Dev.to AI
📰
Who’s Afraid of Chinese Models?
The U.S. should focus on developing open alternatives to Chinese AI models, rather than fearing them, to maintain a competitive edge in the AI landscape.
Stratechery
📰
I compared the real cost of running LLMs on AWS - here's when each option makes sense
Learn when to use each AWS option for running LLMs in production and understand their cost implications
Dev.to · Jerzy Kopaczewski
📰
Building a Character-Level Bigram Language Model from Scratch with PyTorch
Learn to build a basic character-level bigram language model from scratch using PyTorch, understanding the fundamentals of neural language modeling
Dev.to · Mohamed Heni
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →