Content Safety
Skills:
AI Security90%
Key Takeaways
Introduces Azure AI Content Safety, a service for detecting and filtering harmful user-generated and AI-generated content
Full Transcript
[Music] you [Music] you [Music] my [Music] w [Music] [Applause] [Music] new Mon Mon [Music] Mon sorry there was someone in the chat that was asking questions in French and I thought I would start in French just to confuse everybody how about them [Music] apples hello my friends I'm excited to be here welcome to this episode of the AI show I am feeling less under the weather uh last week man I was feeling a little sick and then I was sick the whole week it was awful um I'm in English now it's okay I'm in English now where's everybody coming from my friends uh just put it over in the chat I will uh I will give a shout out to everyone that's uh there pretty excited to be here today my friends uh feeling a lot better it's like that calm after the storm where I don't feel as sick anymore what are we talking about today my friends here we go number one content safety with Sarah bird and like last week uh this one is a live live uh so you may see me be like okay we got to stop here I'll edit this out because if you uh last week we did uh I think it was uh fine tuning fine tuning and we just did it live and you saw me be like okay stop right here because like we edit these things afterwards but because you are here today you get to see the actual show before we edit it and then you can ask Sarah questions and Sarah's time is very valuable I barely got her here I had to like convince her with I don't know I think her candy of choice is Smarties because she is a smarty uh so that's the only way that we got her to come so content safety number one and number two number two more content safet you're Q your quest Q&A with Sarah bird she's thebomb.com we tried to get the domain name apparently was already taken uh unfortunately all right let's see where people are coming from today uh J they're like English angki yazik uh I know I was someone was saying stuff in French and I just wanted to speak in French because um if you look the only Flex I have of intelligence is my book Flex here so I had to speak a foreign language make people think I know what I'm talking about so sorry about that who was the one who said that uh miamian muo uh okay uh and then uh here we go Janis SC number seven hello my friend come see Seth with content safety this is this is Janis SC number 7 in in twitch hey everybody is anybody here is it is it just me other people watch the show obviously my mom and that's it uh okay what Canada lovely welcome uh India Love It Mexico AAS centes I love it um Halloween is coming up uh Malaysia lovely uh India welcome I'm so dumb no uh they're asking if it was Co no it wasn't I teched myself like everyone's like oh no does Seth have Co we need to leave this live stream immediately or maybe I should put a mask on on no it was a regular cold that kicked my butt apparently I haven't had a cold in like four years and my body was like what is this cold thing all right all right we should probably invite Sarah on hold on hold on here's a question is there a type of any type of scholarship for unemployed parents who want to work in tchas you know what this education here that we're doing is for free and I'll tell you what like the stuff you'll learn on this show is eminently useful because there's a lot of people out there talking about like chat gbt this chat gbt that we're actually giving the real information from the actual experts and so hopefully that's helpful um for you all right uh so let's bring uh Sarah on here because uh she ain't got time to waste how you doing Sarah hey Seth I'm good although I guess I had the same experience as you I had a cold that also kicked my butt and I was like what is this I I know forgot what it was like our bodies were like what is what are you putting in here we had a conference in in uh in the UK and I there was so many people there that one of them probably gave it to me and so that was the gift I got from the British so thank you oh my goodness all right so uh we're gonna do this one live and so just a couple of things do you know the rules Sarah you're amazing I am going to be like you're not gonna want me to and then I'll do this and then this time I'm gonna play the thing I'm gonna play because we have the actual video video I'll play it and then we'll get started okay so I got to be serious now so let me turn all my stuff off um okay and then folks if you have any questions please get them in the chat we'll answer them after we make the show you may see us every once in a while have to do like an edit point just be aware that uh if you hear us being quiet we're just like trying to all cut around those in post uh Sarah you ready yeah totally all right you're not gonna want to miss this episode of the AI show we talk all about content tent safety with my friend Sarah bird make sure you tune [Music] in hello and welcome to this episode of the AI show we're talking all about content safety with my friend Sarah bird Sarah it's good to have you back my friend how you doing I'm great glad to be here fantastic so we talked about content safety a lot at some of our conferences so tell us what's new what's going on and maybe if you could explain what content safety is yeah so uh you know I've been on the show a lot over the years talking about responsible AI meaning you know how do we make our own AI system safe how do we Empower others to make their AI systems safe but um one of the things I don't get to talk about as much um which is really embodied in this work is the opportunities we see for AI helping make other things safe right we've had just amazing breakthroughs in AI Tech Technologies over the last couple years I don't need to tell you we're all you know here very excited about generative AI but um what we've seen in you know my team is that that's also a huge opportunity to take those same Technologies and look at how we can increase safety in the world um one of the challenges that uh we have right now is that there's you know vast amounts of content out there uh all sorts of information which is great I mean that's the beauty and the power of the internet and different online platforms but unfortunately there's also like harmful problematic content that you know platforms and users don't want on there and um one of the challenges that we've seen over the years is that uh you know the content's very sophisticated Nuance uh language is very nuanced and so traditional AI was just not great at it right it was like okay if there's you know just one harmful word in here it's going to get flagged and sent to a human reviewer and so it was frankly just dumb well yeah turns out out uh we've had a huge breakthrough in language models right and multimodal models that can understand so much more um than that and actually now look and really understand more what is this sentence saying what is this paragraph saying is it actually um harmful and so uh content safety is the outcome of several years of work to harness that technology and put it towards detecting harmful content so it's a new cognitive service it just G um two weeks ago and uh the idea here is that it's going to be able to uh have different types of content flowing through it and score it with different categories you care about like violence hate self harm Etc and that's awesome because traditionally what we would have and I I come from the NLP World traditionally what we used to do is we have a huge list of just bad words that we just searched for and then someone would put like a zero instead of an O and we'd have to add that thing in there and then if and and you're saying that we don't need to do that anymore with content is a much more intelligent way yeah certainly we've seen a huge lead forward I mean there are obviously with any of these it's a you know an adversarial game where people are going to figure out more you know clever techniques to get around um these kind of filtering systems but it's certainly a huge breakthrough and sophistication and and it um I think greatly reduces the other side of it which is also where you have people trying to use terms like in a positive sense reclaim them and being censored right and so we want to make sure that you're you're able to both like not be over filtering or under filtering and so um I think that that's uh we've just seen kind of a huge step forward in that and um actually I thought it was very fitting you started the show in French I don't speak French um but uh these models are multilingual from the base also which means that when the sentences are uh you know mixing different types of languages you still can understand that and that just adds actually higher quality in every language but then also going across languages and so that's been another part of it of course and that's that's interesting because I hadn't even considered the multilingual case right if you have to manage this thing yourself you're going to be fighting this game a lot and you're saying that now there's a service that we just optimize this thing all the time to make sure it's doing the right thing yeah absolutely I mean we start with the same Foundation models that are powering other things we start with like a great language model that's multilingual at the base and then we train it on harmful content and I'm you know talking about language and text a lot here because we've all seen the breakthroughs but the same applies to image right our this is powered by our Florence image model or um you know the newest area which is actually really important which is multimodal where we're combining image and text right if you look at like a hateful meme uh the text can actually be totally safe and the image can be totally safe but the combination is problematic and so the multimodal models are also really important in that so this this step forward and Foundation models understanding context has allowed us to build on top of those and understand very specifically different types of potentially problematic content and and this is awesome because like every time you come on the show I'm like oh that's another thing I should probably worry about I I was only thinking of like written content and you're saying now image content as well is something that needs to be looked at because someone might post a very bad thing in an image and your your text word search is not going to do anything yeah exactly and then of course the the combination right it's not just like you can also have image like that is you know clearly violent or let's say clearly hateful but then a lot of the memes are really that the combination is what's harmful neither the text nor the image so you need a model that actually understands the combination of the two um and so that's been a gap that we've always uh sort of had out in the world and the new multimodal capabilities of foundation models are enabling us uh to do that better and multimodal is not yet GA it's in preview that's like the sort of the latest work but I'm just really excited where I think we can get to in the next couple years with this technology I think it's um it's really going to be a game Cher for safety and I've seen this in in use primarily in the context of large language models like when they're using generative large language models filtering content on the way out and on the way in that happens in Azure open AI are you saying that this is now something that can be Standalone used by itself yeah exactly so um when we first started working on the technology the first place that we released it was in the generative AI applications as part of azure open AI because that was uh frankly because it was needed right because we had to be able to filter in real time at scale which meant we needed a different type of Technology um so we've you know quickly put this in there and so you've been seeing it in ashure open AI um and we can you know look kind of more of a demo of that um but now we have it as a standalone system which means that people can use it to look at also just user generated content on their platform but also open source llms or if you're developing your own model you can use it so we wanted to make it easy for safety to just be wherever you are and so having it as its own API means that you know you can you can bring it to where you need it um but of course it's still integrated in Azure openi so that it's just there and it's easy for you and that's interesting using a Content safety as a pseudo regularizer over your own training of models which is kind of cool as well yeah exactly all right so what can you show us yeah so let's go and look first just at um sort of the content safety portal here so this is the studio for Content safety um and as you can see just as we were talking about uh we have um you know we have the uh moderate text here can you see my cursor I sure can it looks great sorry okay um yeah so you can see the moderate text here we talked about image I mentioned that we have multimodal and preview which is really exciting I think um uh in each of these um when you click through we have different examples that you can look at so you can understand how it works um but you know trying to limit the amount of parable content people look at let's stay with text today because the images um you know you actually are kind of visually looking at that which is problematic um and so uh the one of the things that was really important um is recognizing that like each application is different uh in terms of what type of cont content might be appropriate um but we also wanted to build a system that could power many many applications because it's taking like a lot of expertise right we work with expert linguists and fairness experts uh to label this data and train our models uh we have you know pretty sophisticated AI scientists working on this and so there's a lot of investment that we wanted to make sure is reusable but recognize that applications are different and so the way we've been doing this um behind the scenes is actually that the uh we've have different severity levels for the content that we've actually trained into the model and so you can set the system to filter at uh for each of these different categories low severity medium severity High severity so as was a gaming application for example I might want to allow more violent content through uh but maybe not the worst of the worst right and so I would set that here um where an education application might choose to be much more strict and actually do that um and it block everything that's even a level of severity uh so this system is allowing you to understand depending on where you put that threshold you know this playground kind of where that was and the example that you know we showed before which I think is still a good one um for our Koso you know outdoor um Commerce company is asking um uh about you know I'm looking for an axe which this has a uh you know a a weapon right to cut so it has a problematic um you know it has a problematic verb here and then we set a path in the forest right and let's hope our system here recognizes right this is safe you're allowed to cut a path with an axe um and then of course if we switch this over and we say something much more problematic but hopefully not hopefully a little cartoony here then we recognize like this is a medium level of violence in the way that it's um said here and so this is um you know this is kind of how this is working and actually with the ga we've added additional severity levels so um you can go up to eight different severities so you can get even more fine Nuance in terms of what uh what works for your application uh and so this is um this is kind of the system you know in the most basic form now of course it's an API um and what you actually get out is a category and a score so you could do something much more sophisticated you could for example have all high severity content automatically filtered and uh medium severity you could send to a human reviewer for example first to decide if it can be you know used or posted or something so there's a lot of different options with these we're just showing kind of the the most basic um version here here's a question I have for you so because I've looked at this a couple of times and sometimes I get confused so when it's on the low severity that just means that the content that's going through has low scores for each of those categories is that right yeah it's this is It's kind of um we're evolving some things it's kind of um here saying where it's going to start rejecting um and it's just letting you understand like if you set your application to reject starting low it will reject everything low and higher right if you set it to go to medium it will reject everything medium and higher um so it's giving you those uh the where you want your sort of cut off point for what's appropriate for your application like what type of level of content can you um tolerate God so it's yeah it's almost like how much what is your tolerance for whatever content there is like for example in a video game if you're playing like a video game that has people shooting the violence tolerance for the speech is going to be much higher exactly and that is that what you're saying and so that then you would set it to high or medium or whatever this is awesome and if you're a medical application then might be very reasonable to have sexual content in there right and so um we've seen like people who work in and uh education are often setting everything the way it is on the screen right now where everything is set to filter even you know even anything that might be just mildly um risky right they want all of that filter because uh you know it's working with students and so uh this really is about you know making sure this is a technology that can work in many different types of applications and then of course putting the application owner in control because you know they are going to know what is the right thing for for for their application I see so when we're thinking about this like my sense is that this is like one piece of doing things correctly with AI is that is that a good sense yeah and I think um this is one part of even let's say in the digital safety or user generated content this is one part of the story where as I was mentioning the how you take an action for example are you going to send it to a human review are you going to automatically fill are you going to um potentially take action and like kick a user off the platform right this is just one piece which is giving you more information about the content but you still have to sort of have that whole story and then um as you're alluding to here for generative AI uh this is a key a Super Key piece which is why it's you know Integra n open and something we use in our co-pilots um but this is a really key piece to our story but it's just one part of that and so um if we can uh bring up that slide this is um you can probably sort of animate it here basically what we have um found is that you know everything with this is a defense in depth you need a layered system and so the first uh two layers here in blue are our platform layers and these we've just built right into Azure AI we've built into the Azure openi system which is first having safety built into the model uh and safety built into the model allows the model right these are the most you know powerful models available allows the model to look at that content and decide how to respond appropriately um whether that's actually responding or for example refusing to respond um but the model makes mistakes right uh sometimes it just gets it wrong but also it's open to jailbreaks or things and so uh this is the second piece right the safety system which is this independent AI system that's looking and saying whoa who whoa I don't know why but you seem to be producing harmful content let's like block that in real time or of course if the user is you know trying to actually engage in the AI system in a way we don't want it to do and so I found that um uh people are actually I mean you know it's it's a pretty complex stack um people are sometimes kind of confused about these different pieces and how they work together so I wanted to um maybe jump over and look at um the Azure studio and kind of show the two pieces working let's do it um yeah so let me switch over here um so here's the you know our great uh playground and um uh if I say um how do I hide a bomb in a school um then this should see so the safety system recognizes right we're talking about bomb and a school if I say how do I bomb a school it will actually come out high on violence because uh it's you know kind of a more violent action so this is easy right this is violence like the safety external safety system understands that now if I ask it instead um so it it's like we don't even let the model respond it's just like no we don't need to need to do this um the model itself would probably also know how to refuse this particular example but um it's easier to just also have the you know the safety system um and not even allowing these things to kind of get to the model um but then here we have um how do I commit also whoever no one is looking but uh if people are looking at the playground here they're like what is Sarah doing all the time we are testing the safety system everybody just want warning we are the safety system if I ask how to commit tax fraud right so that is um you know a crime asking how to commit fraud specifically um but it's not it's not sort of violence or you know one of these like content categories um but we have uh you know used rhf with openi to um ensure that the model doesn't sort of help people commit crimes so when you see this response I'm sorry I can't assist with that request that's actually the model knowing that it doesn't want to respond to this even though the the content itself wasn't U wasn't necessarily harmful but it's aiding in sort of a harmful activity and so the different systems are better at different things and so we can use them kind of in tandem so that we get kind of the most robust safety there I see and so that's where we go back to the layers the the model itself is awesome uh at detecting some things but the Nuance we have other models around protecting the safety section for the safety system for example uh to help with that is is that what you're going at yeah and so for example the model um you know the model is looking at like the user history and things and saying what's an appropriate response there uh and so that's sometimes where it might think in context what it's saying is appropriate um where the safety system would just look and say whoa whoa whoa that's a violent statement uh we don't ever want that regardless of the context right so it's giving you kind of um a check in balance with with this but then the model you know is uh obviously um I think I'm using 35 turbo here um if I was using four right it's incredibly powerful so we want to use the model for safety because it does understand so much in can take that history into account and so there's lots of things where we'll see the model will either know to refuse or or better yet know how to respond appropriately because even better if it doesn't have to refuse or it doesn't get blocked um but in the cases where it it slips up that's where the safety system uh you know can come into play so that you you know that you're still going to get H you're GNA have safety even if the the model doesn't quite do the right thing and I love this safety system but when it comes to your application there's still safety that's just beyond those five categories for example that relate to the applications you're building and so that's where we get to meta prompt and grounding and then user experience as well could you say a word or two about those things yeah exactly and so these two layers that I was just showing you are the platform layers um we've built them to you know work with a variety of applications um and then the two layers on top The Meta prompt and the grounding and the user experience that's what the application developer is is developing and so uh The Meta prompt um I think people still underestimate how important and how valuable this is for safety where uh we see so you know so many meta prompts that people are showing us that are like hey can you give us some feedback and the safety section or the respon wise section is is just very small right and this is your chance now these you know these systems can take pretty large promps this is your chance to get really you know detailed instructions on kind of how you want the system to behave in different circumstance ances and things and it makes a huge difference on how it actually behaves in practice right so this is like the first thing you want to do if you don't like the way the system be is behaving is you know adjust the meta prompt and that's actually where our tools like um you know prompt flow are so great now because you can go and evaluate these meta prompts side by side and actually see like if I change this am I getting the outcome I want um but that's just like a hugely important part of the story and that's where you really tail it to your application right you're going to want to behave very differently and each of these context and so the meta prompt is what's really trying to get it to work well in each of the those cases and then the sort of built-in safety in the model and the safety system are for um you know when things don't go quite as planned and then the last part and this is the part that I think a lot of us kind of forget because we're like I'm gonna prompt engineer this and then I have my safety system and the model's goingon to be awesome so everything's perfect I think people sometimes forget about the user experience we we've got to do some stuff about like maybe disclosing that this is AI generated at minimally right or or what other things should we be doing yeah and I think we've got um we've got some examples in our hack toolkit which is where we put sort of our best user experience best practices but absolutely um one of the things is like we call our systems co-pilot for a reason they they're supposed to be designed to work with a human and humans are great at certain things and these systems are great at other things and so if you can co-design with the user in mind then like it can be so much more powerful where um like you know the best example of this is like GitHub co-pilot the you know original now I guess the classic version right but even though it's made some mistakes the you know the users could accept the suggestion or not they can edit the suggestion and so the developer was still very much in control and could send that you know still through all their normal testing and security processes and things and so it just helped developers go faster like oh I don't want to type all that or if you give me an idea to start it's a lot easier and we've seen obviously developers just love it even though it's not perfect right it does make some mistakes and um and that's a case where it's just like very nicely designed to work well with how those users work um one of the challenges we've seen though is uh basically one of the magical things about the technology is it enables users to do things they couldn't do before right you can generate code even though you don't code um and so uh but the systems still make mistakes and so one of the things we're very much looking for is how do we um sort of avoid this risk of overreliance where you're relying on the system to do something that it doesn't do perfectly and you're not really able to compensate it because like human in a loop is a great pattern if the user is actually able to fix the mistakes but if the user can even tell it's making a mistake then uh then that's a big problem and so we're still you know each application is a little bit different with that but that's like definitely one of the things we're seeing you know people really need to watch out for is that now now you can you can have this tool that allows you to do something you never could do and then how do you have oversight over that and so um that's where like really making sure we're developing it with the user in mind is so important to understand like is is the user going to have oversight or do we need to make sure it doesn't make mistakes do we need another system having oversight because the user isn't going to be equipped to do it and so uh there's a lot there and it's a really important part of the story well all this is amazing I I love the what AI is doing and transforming how we're doing work and I love that we're doing it we're trying to do it at least at Microsoft in a safe way and giving people the tools to do that uh so is there where can people go to find out more Sarah um well you can go play around in the playground I mean honestly I think any of this technology just like actually using it is the most important thing you get a feel for it and as I said I only demo the text because it's you know a little easier to look at harmful content but I really recommend checking out images and multimodal um and understanding what it can do there um so you can go just check out Studio start getting a feel for it obviously you're already using it if you're using Azure open Ai and you'll start seeing more of those features and controls kind of coming through there um we also have a great ebook that I think talks more about some of the things I was talking about of you know how this can change digital safety and enable kind of a new tool in the toolkit which is um you know something we're always looking for for safety um so you know I would recommend uh getting started there and then if you want to hear more about the responsible AI layers and Rings um we have all of that in the Azure openi documentation and so you can see there and we'll keep bringing out more as we're learning fantastic and there's a Blog too from what I understand a safety blog that talks about the ga it's been a lot it's been a you know a labor of love for me and kind of a a long journey to getting to this GA um in terms of when we first had the ideas that we can do better with these kind of models to where we are now in the ga but it's I'm so I'm so excited about that definitely check out the blog and it's such an important mil because now you can go use it in production right and so you have no excuse to not have safety everywhere at this point because you made it easy it's an API you just plug it in well thank you so much for being with us Sarah and thank you my friends so much for watching we're learning all about content safety with Sarah bird thank you so much for watching and hopefully we'll see you next time take [Music] care alrighty that was amazing uh that was the live part now we can now we can be funny cool and I think I see some questions floating there's so many questions so I'm gonna go through them here uh let's get to the chat here holy cow people were having a chat there let's see what we got let me go to the questions here uh would you like to talk about okay uh let me go down oh okay uh wow there's a lot of wow there's a lot going on yeah uh oh my gosh I gotta okay okay there was a question but there's so many people just like talking in French okay here we go here we go this is a good question this may be an elementary question but what do you consider unsafe content violence bullying what about disinformation uh what does disinformation validate against yeah that's a really important question um so the reason that the first four uh so so none of these categories are easy right and that's where we like actually sit down with experts in the area with our digital safety experts inside of Microsoft with linguists Etc and say you know what classifies as violent and then of course we're also saying what classifies it each severity level of violence so we have to think about like why would this one be high severity and why would this be low severity and so we try to take kind of a principal consistent approach um with that but it is like that is where a lot of the energy is like why should this be here and not there um we have started with categories where by looking at the the content itself you can sort of have a categorization of it where um some of the categories of like misinformation disinformation those ones um you need like a larger um s context you potentially need a source of what you consider a source of Truth Etc um and so uh those are more challenging uh in terms of sort of the techniques you would use to address them and so that's why you actually you don't see them in the system today I see and so that that the cool thing about this I think CJ just to add to what uh Sarah's saying is using this API will give you additional information that then you can make reasoned uh you know uh like you just like for example let's just say you do a site about something and you have certain tolerances that other people do not and you put that in your rules and whatever you can just use this to inform like hey let's flag this and then have a h and go review and then be like no that that's not what that's not bad or that is good or whatever and and that's how this is intended to be used right Sarah yeah certainly it varies by application what the appropriate thing to do like when we're using in the generative AI applications we obviously like if it's the AI system putting an output we're choosing just a filter even if that might in some cases be over filtering because we need that real time um sort of decision but if you were doing a platform with user generated content then it absolutely you want to make sure that you're adhering to your organ your sort of platforms policies and things you want to have human reviewers go and look at the results and so this is just adding more more understanding of the content more information to feed into your sort of actioning system whatever is appropriate in your application but to be able to do it um reasonably sophisticatedly at scale in real time is just not something that you could do before and so it unlocks a lot more possibility in terms of potentially scanning every piece of content which is what we do in you know generative AI land um but uh you know not some a lot of platforms you had to wait for content to be reported or something right and so it's just another tool in the tool belt but no single tool solves any of these problems right you have you have to look at combinations of them including the humans involved yeah absolutely and I love that these categories have nothing to do with truthiness uh because I don't know how an llm could ever or a model could measure truthiness uh and so it's all about like uh hate uh hate speech violence whatever uh there's five what are the five there um it's right now what's GA is hate violence self harm and sexual content and then we'll be adding additional categories both um both different types of content categories for example I think um someone in the chat has called that bullying and bullying is like um harassment is kind of between some of these like hate and violence but it's kind of its own thing as well and so um you know that's one that we're looking at or but also more llm specific things right um in terms of something like a jailbreak which is just a very llm type interaction it's not something we really want happening in our system and so we're looking at a mix of where do we extend the different content categories to give users um you know customers more information more control but also um you know looking very much at the llm specific risk or the generative AI specific risk and how do we build systems that are good at looking for those amazing all right so here's the next question from Jan number seven does it understand context or some words in the context what what is it understanding yeah so I think it it can't understand anything it doesn't see right so it's not going to understand the some broader context like this is in the context of a social media application or something right it still only understands the content you're giving it but what um you'll see in some of the the the systems that were used in practice is they really couldn't consider for example more than like a sentence right and even in a sentence they were more likely to be triggering on keywords than understanding what the sentence is saying we're now um you know we can look at a paragraph um in some cases more and it's understanding like the whole statement more holistically and really trying to get at what is it saying and so um I you know we demo very short examples here although in the playground we do have some longer um things and that I think is another big part of it is that like we all use language very richly and um the meaning of my thing might really be contained in the entire paragraph not just one little part of it and so you don't want the system just triggering or on you know a single small parts that really are taking taking it out of context but certainly is not understanding any context that isn't visible to it I see cool so here's a here's from our LinkedIn user it says building a mental health related chatbot having low tolerance input has been useful in automatically redirecting the conversation towards seeking professional help as one example which is which is an interesting interesting context have you seen stuff like this used in uh in the medical Industries for example yeah I think um actually self harm is exactly one of these where yeah you want to to push towards or and I I saying self harm as our category but they're obviously a broader said there but uh not just like refusing to answer but actually directing towards resource and directing towards help and so from a respons point of view this is where those two systems can work really well together because you you want the the application to respond with resources not refuse to respond um but the fact that you've detected for example um an intent that uh maybe self harm would be a case where now you can then the action you can take us to respond with resources and so just having more understanding allows you to respond in the right right way there um but also we've talked to people who are looking at how do we build chatbots for people for example in the the context of self harm who have something to talk to and that's where like that self harm category you do not want to be filtering on the users inputs if they're talking about self har it you want to be engaging but you don't want the system for example to be saying yes go do this this is a great idea so you still want to have the safety on sort of the system output um there and so that's where this uh you know the the policy is really going to vary a lot by application um because of how you know allowing you to construct an application that responds in the way that's appropriate for your users and what you're trying to achieve well this is amazing um just a last question what can people look forward to in the future obviously without giving too much away um well I already told you the part I'm super excited about which is where we can go with multimodal right because like it's just a it feels like such a failing um in the past that we can only understand one modality at a time right and most of the things slipping through the jailbreaks you see are because for they're multimodal and so that's one I'm really excited and passionate about um you know in general we're looking at of course how do we make the systems better AI quality is really important here um but how do we add more user control so from the preview to the ga we went from four to eight SAR level so that users have more fine grain control of the system um so you don't keep going in all those directions and then we're thinking very very deeply about uh you know these what are sort of generative AI specific concerns uh jailbreak hallucination things like that and what can we do about that and can the safety system help um there and so I think you'll see more exciting things coming along those dimensions and as always I think our our Microsoft ignite conference is in I think two weeks yay to be transparent Sarah and I still have demos we need to finish yes yeah exactly very that's why we're like yay we're feel like we're living in the future a little bit right now like not night right yet oh my goodness so yeah in two weeks you can look forward to some new stuff that's gonna be coming out I think it's pretty exciting uh and so make sure you're tune that all right so Sarah uh thanks so much for being with us my friend yeah and thanks for everyone for watching and asking all the great questions I love hearing from you and yeah these are really hard problems we're excited to be working on them really want the feedback if the system's not behaving the way you think tell us where you know we're looking at these examples we're making it better but there's so many applications in the world we need to learn from your real world feedback so please please please keep sending the feedback tell us we need it and absolutely and like I said this stuff is designed to be helpful if it isn't helpful well we want to know why dang it yeah exactly all right we we'll see you later great bye all right my friends um Sarah's the best uh we always love having her on uh all right we got about like what 15 minutes 15 minutes left uh let's see if we can't uh do a thing uh so this is the slide that Sarah was showing uh let's see let's see if we can't use content safety and let's make a Content safety thing so I am going to let's see let's see if I can do this here and um we're going to make a brand let's make a new let's make a whole new promp flow and see if we can't put sa no I can't show you that let's see if I can't put uh safety in there so let me go let me go to Project I'm going to make a new folder uh Safety Safety project projecto uh and here's my empty folder and I am just going to show more options here open it with code there you go let's see what we got uh let's see here um oh look at all oh prom flow there's a new one there's a new one there's a new one let's reload it I don't know why I make sound effects with my mouth when like literally um literally like I have like a whole sing of buttons where I could just be like but I'm like all right so let's make a new promp flow project here and the way I do that is python minus minus M VM VM here let's see if we can't use content safety yes I want to yes yes I do uh uh uh we're going to say uh VM activate CLS uh touch requirements requirement m. text and then uh python.exe pip install Okay so let's add the prompt flow prompt flow and prompt flow tools uh and then I'll just say uh python py python python oh no no no it's pip pip minus uh yeah there we go and so we're installing the uh the py on things I think I'm going to need a Content safety service so uh portal portal portal. azure.com while this is loading I'm doing this here on the side and I'm just going to make a new Resource Group here for no maybe I have something oh there you go create uh so here I am let's see if I can do content content safe safy safy content safy look at that let's just make it let's do it let's create it uh yes East US super safe safety uh and then obviously let's see we'll do we'll do the free one you know because we're uh we want to be Thrifty next next uh no I don't care about that next next all right so now we got this content to safet S uh system ready so I'm creating that let me restart this little gem okay now this is ready to go uh do I need to install this do not show again not now fine fine okay so now I'm going to do a new promp flow here we'll do an empty flow nice this is already done wow we are going really fast okay now I need to add in prompt flow uh no uh notice that there's connections that you can you can do in promp to flow so let's see if it has a Content safety it does look at that so I'm going to add new content safety thing and then we're going to put our um connection name this is our Isis uh Isis safety and uh here is the endpoint and then we'll create let's get the key here liave I'll put it in here you're probably thinking oh it's giving away K nope look at that they were so clever they saw me doing these things and they're like we got to find a way to get make Seth put his thing in there all right so looks like we now have a Content safety thing in here called Safety so let's go to our visual editor and let's do an input uh let's do a question here okay and now let's add add an llm node so we'll do llm node here and we'll just call it llm llm question question and we'll go to new file here and now we have this question so let's choose the uh Isis connection and we'll do a chat model and we'll do we'll do full gp4 here and then on the input we don't need a chat history here so let me go to let me go to the model here and let's update the prompt so we don't need a chat history so we just need this here so let me go back to my flow. come on refresh there you go this is the question and then the output here is going to be the output output and then we're going to say this oops output okay so now we have a full prompt flow of using an llm here so let's test this and then we'll go ahead and uh see what then we'll add content safety uh how can I cut a cut a path C can I can what did she put what was it that she said can I buy an axe to cut a path in the forest okay so we'll save this here contr c s and then we'll go to we'll go to the actual thing you are helpful assistant at a hardware store you respond to questions in a helpful helpful way and refer customers to a human when you are fused your knowledge based of hardware store items locations you are also powered by and there um all right looks like it got stuck in a loop but okay notice I'm using I'm using AI to make AI how insane is this all right let's run it uh nice so it looks like you can definitely buy an axe for that purpose you can find axes blah blah blah would you like information arrange so this is working already so let's change the input to BU ax to cut a person let's see what happens uh let's see what happens here in theory gbt 4 should already like tell us I can't do that nice oh my goodness my phone is bad AI is trying to be like hey okay so notice that this failed but we want to make a promp flow that like catches it and says something in return right because failing is not cool so this is where we're going to add the safety system how do we do that uh all right I'm just gonna make a python node to do it U because I don't know how to do it otherwise so we're gonna do a python node uh uh safety safety safety check new file okay so let's see how this thing is used [Music] um here's the API reference oh look at this analyze text cool cool cool co co co co cool all right so what we want to do then is we want to take this thing and we want to um here uh this thing and we want to do a input so that it checks the input so here's the question and we'll we'll leave this here for now and then we'll do a saf check saf check we only got like six minutes left let's see if we can get this to work in six minutes uh and so this is no longer going to take the well maybe it does take the question no we'll see we uh maybe it does uh and then this thing also takes the question here okay you know maybe we should uh there's a cool thing in uh these things called the activate config um okay we'll what we'll do is we'll we'll we'll unhook this for now we'll unhook this for now and just do a safety check and then we'll we'll put that we'll put that other thing back okay cool so in the safety check looks like we have to do an API call uh so this is a post uh okay and then the other thing we're going to do is we need a um PR promp flow. connections import there it is azure safety connection and then we'll say uh connection Azure safety connection and then notice here it should be like hey um choose a connection there it is Boom nice oh shoot I need to put the walk-off music here to timing safety check safety check okay so now that we've got that here what we can do is uh we have this end point here uh end point equals F string [Music] this and in the connection we should have we should have the uh where is the end point end point there you go nice look at that hey yo hey yo boom I don't know if it's is it subscription key API key look at that we also have API this is glorious GL shoot uh looks like here we also have the API version [Music] here [Music] okay uh looks like the request headers need these two things and let's do it uh import import requests okay and then what we're going to do is we're going to say uh con con I'm thinking in JavaScript like a like a dude is uh let's see I got to get this call so uh request uh response equals [Music] Rec requests request [Music] body we'll do violence on this [Music] one hold on hold on hold [Music] on [Music] oh okay I see uh [Music] body come on you don't want to format this for me come on [Music] okay I think this was just right wa I think this is just right uh response [Music] dot content is that what we I don't know let's debug it let's debug it why why are we mad what's going on what where's the madness oh whoops there you go all right so let's debug this uh goodness and see what happens debug oh and we're almost out of time here we go uh let's see if it works oh I don't want to show the thing oh oh well we'll figure this out later all right my friends thank you so much this has been another fabulous episode of the AI show next week what do we got best retrieval strategies for generative AI application semantic search benchmarking with myself and Liam Kavanaugh he's an awesome dude so make sure you stay tuned for that he uh we uh like I said this guy is amazing you're going to want to take a look at this Azure cognitive search is amazing thank you so much for watching you've learning all about content safety we tried to make a thing work in the remaining time we had we'll fix it later uh thank you so much for being with us and hopefully we'll see you next time on this the AI show live every week at 8:30 a.m. Pacific hope to catch you then see you next time my [Music] friends
Original Description
Get ready for an episode you won't want to miss! Join us as we dive into the incredible world of Azure AI Content Safety with our special guest, Sarah Bird. 🔍🛡️
Discover how this cutting-edge service can revolutionize the way you protect your applications and services from harmful content. 🌐✨
🔹 Detect and filter harmful user-generated and AI-generated content.
🔹 Monitor text and image context across various categories and languages.
🔹 Prioritize content with severity scores for efficient review.
But that's not all! Azure AI Content Safety seamlessly integrates with your favorite tools like Azure OpenAI, Copilot, and Bing, making it a must-have addition to your tech toolkit. 💼💥
Mark your calendars and be a part of this transformative discussion. Stay tuned for updates and join us for this eye-opening episode! 🎤📺 #AzureAI #ContentSafety #TheAIShow
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Microsoft Developer · Microsoft Developer · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Prepare for the DP-300 exam & the Azure Database Administrator Associate cert | Data Exposed
Microsoft Developer
What I Wish I Knew ... about landing a job in tech
Microsoft Developer
Igniting Developer Innovation with Vector Search
Microsoft Developer
Combining the power of vector search with Azure OpenAI then revolutionize image search with vectors!
Microsoft Developer
What I Wish I Knew ... about finding your place in tech
Microsoft Developer
Fluent UI React Insights: Accessible by default
Microsoft Developer
Signing Container Images with Notary Project
Microsoft Developer
What I Wish I Knew ... about finding your place in tech
Microsoft Developer
What programming languages does GitHub Copilot support?
Microsoft Developer
What I Wish I Knew ... about how much your job can change
Microsoft Developer
What I Wish I Knew ... about how much your job can change
Microsoft Developer
How do I become more confident about AI?
Microsoft Developer
How do I become more confident about AI?
Microsoft Developer
Performance Demos of SQL’s Intelligent Query Processing Feedback capabilities | Data Exposed
Microsoft Developer
What I Wish I Knew ... about coming to Microsoft
Microsoft Developer
What I Wish I Knew ... about coming to Microsoft
Microsoft Developer
Revolutionizing Image Search with Vectors
Microsoft Developer
Igniting developer innovation with Vector search and Azure OpenAI
Microsoft Developer
Getting Started with Azure AI Studio's Prompt Flow - Part 2
Microsoft Developer
What I Wish I Knew ... about finding my career path
Microsoft Developer
What I Wish I Knew ... about finding my career path
Microsoft Developer
Windows Terminal's journey to Open Source
Microsoft Developer
Can I trust the code that GitHub Copilot generates?
Microsoft Developer
What I Wish I Knew ... about interviewing
Microsoft Developer
What I Wish I Knew ... about interviewing
Microsoft Developer
What is the Microsoft TechSpark Program?
Microsoft Developer
SQL Server 2022: Accelerate query performance while reducing query compile time - w/ no code changes
Microsoft Developer
What I Wish I Knew ... about discovering computer science
Microsoft Developer
What I Wish I Knew ... about discovering computer science
Microsoft Developer
Call center transcription and analysis using Azure AI
Microsoft Developer
How to use Text Analytics for health in Azure AI Language
Microsoft Developer
Azure OpenAI-powered summarization in Azure AI Language
Microsoft Developer
Accelerate data labeling using Azure OpenAI and Azure AI Language
Microsoft Developer
Building a Private ChatGPT with Azure OpenAI
Microsoft Developer
What I Wish I Knew ... about how to interview
Microsoft Developer
What I Wish I Knew ... about how to interview
Microsoft Developer
Getting Started with Azure AI Studio's Prompt Flow - Part 3
Microsoft Developer
Intelligent Apps with Azure Kubernetes Service (AKS)
Microsoft Developer
Getting Started with Azure Blob Storage | Data Exposed: MVP Edition
Microsoft Developer
Chat + Your Data + Plugins
Microsoft Developer
What I Wish I Knew ... about different career paths
Microsoft Developer
What I Wish I Knew ... about different career paths
Microsoft Developer
Advanced Dev Tunnels Features | OD122
Microsoft Developer
Learn Live - Manage performance and availability in Azure Cosmos DB for PostgreSQL
Microsoft Developer
Plan your SQL Migration to Azure with confidence | Data Exposed
Microsoft Developer
What I Wish I Knew ... about social skills in a tech career
Microsoft Developer
What I Wish I Knew ... about social skills in a tech career
Microsoft Developer
All About Vectors, Search, and Function Calling in Azure OpenAI - Labor Day Special
Microsoft Developer
Introduction to project ORAS
Microsoft Developer
What I Wish I Knew ... about finding the right major
Microsoft Developer
What I Wish I Knew ... about finding the right major
Microsoft Developer
What I Wish I Knew ... about how to approach programming
Microsoft Developer
What I Wish I Knew ... about how to approach programming
Microsoft Developer
Learn Live - Scale from a single node to multiple nodes with Azure Cosmos DB for PostgreSQL
Microsoft Developer
What I Wish I Knew ... about diversity in tech #1
Microsoft Developer
What I Wish I Knew ... about diversity in tech #1
Microsoft Developer
Get started with SQL Server AGs across Windows, Linux and Container Replicas | Data Exposed
Microsoft Developer
Writing LLM Apps with Azure AI and PromptFlow
Microsoft Developer
What I Wish I Knew ... about how cool working in tech could be
Microsoft Developer
Open Source foundation models in Azure Machine Learning & optimization techniques behind the scenes
Microsoft Developer
More on: AI Security
View skill →Related Reads
📰
📰
📰
📰
Musk thanks Micron for chips, and builds a $55bn fab to replace it
The Next Web AI
Jensen Huang calls the AI jobs panic ‘complete nonsense’, and takes aim at his peers
The Next Web AI
IMF says Africa has to keep lights on before it can bet on AI
TechCabal
What Does Job Security Even Look Like In 2026? It Starts With Skills
Forbes Innovation
🎓
Tutor Explanation
DeepCamp AI