Culturally Aware Machines: Why and when are they useful?

Microsoft Research · Beginner ·🧠 Large Language Models ·1y ago

Key Takeaways

The video discusses the importance of culturally aware machines, particularly in the context of language models, and explores the challenges and limitations of current approaches to cultural awareness in AI research, highlighting the need for more nuanced understanding of culture in LLMs, with tools such as GBT 4 and GB4 being utilized.

Full Transcript

so um it gives me great pleasure to welcome monit uh for the aan society talk series um monit is right now a professor at uh mbz U AI That's the University of AI in Abu Dhabi um but for the longest time he's uh been with Microsoft research uh India and a very um close colleague and friend of mine we we've practically done hundreds of projects together uh while he was at MSR India and we continue to collaborate even now his research interests center around the convergence of language technology and society and uh he's very interested in learning more about um the misrepresentation of linguistic and cultural diversity um in and by foundational models he's actually um an NLP uh scientist and researcher um by training but um over the uh years he's um kind of become much more interested and how language is used and how language is a part of U the social cultural system of any community and how does that impact technology and how does technology impact that um so with that uh I'll hand it over to mon to start thank you kalika and uh good morning good evening to everybody uh it it really feels rather weird to be introduced as an outsider uh to uh to MSR uh because as Kika said I pretty much spent uh most of my you know career in Microsoft so from 2007 to 2023 I was in Microsoft and most of the time I was at Microsoft research only the last one and a half years I was uh in uh Microsoft touring but nevertheless uh the team the photo that you can see with kalika suda and this is from uh back from I think 2014 or so uh when we were working uh in project malange on code mixing um and uh pretty much much uh the whole group that we set up had set up there had been working on uh low resource languages and multi issues of multilingual societies and so on and just a year ago I joined uh mbz UI which is the first University uh on AI uh in the world and it's in Abu Dhabi so that uh for on the the top that you can see is our team at emnlp so we have an entire department on natural language processing because uh the universities on AI so naturally the Departments are like NLP computer vision machine learning and so on and um we have 18 faculty only working on NLP uh so with that uh let me tell you a bit more before getting into the specifics of the talk uh let me tell bit more about my research interests and research Vision uh so as Kika already alluded to uh very U broadly speaking I'm interested in Equitable AI or making AI useful and safe for every let's say language every culture and essentially every user um so uh today's talk will be mostly about uh the culture part which is something I started looking looking at uh in the last two years uh actually we started working on culture back in 2017 uh we had a project at MSR at that time called um artificial social intelligence but I would say we were quite ahead of our time because at that time uh one of uh my colleagues from MSR had commented uh after one after a talk I had given on ASI that you know M Ms can't even do simple Mantic processing and you are talking of complex pragmatic understanding you know first let your NLP models understand simple statements so if and and the world changed in 2018 with BT and Transformer models and after that it has been phenomenal right so I think we were a bit ahead of our time when we started ASI project and then we stopped it at the time uh but but yeah now now again I think it's it's the right time to continue it so it's very exciting uh so um and and I'm not going to talk about language uh in in this particular talk but I I do work on low resource languages as well uh apart from how to make AI useful and safe for everybody the other I'm also excited by the other aspect of how AI can help uh understand culture better uh and uh I again I think AI has a strong potential of revolutionizing uh and and changing the way we do social sciences and cultural studies uh we already are seeing the early signs of it and uh I have a uh some work already going on there as well but again I won't get time to talk about that uh so about culture uh of course uh languages we have done a lot of work at Microsoft showing how llms under represent misrepresent most of the world's languages uh something like 88% of the world's languages spoken by 1.2 billion speakers are not at all represented in any of the generative AI models uh but recently uh we did a similar study for uh AI uh music models so music language models to be specific and we found uh a very similar statistics that the music of the global South is very very uh under represented of all the data sets that we analyzed together in terms of number of hours uh Global South is represented in 5.6% uh of the data set and remaining is the global North and and if you do analysis across many other dimensions uh you will see uh similar skew within music I'm saying now music is one of the fundamental monu if you could explain what a music uh model is music language model is a music language model is uh a model that generates music uh given uh description in language so it is like imagine di where you describe an image and uh in language and it generates uh the image here instead of an image it generates some music so something like generate uh a happy peace on violin in Boven style and then it generates a piece or or you could get as elaborate and Detail in your prompt As You Wish now the happy piece on violin by Boven it would generate quite well but if I had said generate a poignant piece in ragab bavi in sitar it would be very very bad at that so that's what we show in this work uh that uh Global South uh music genres from Global South are very underrepresented in these models um so so that was just one um um snapshot but you can see this theme coming up again and again in many aspects and there's a big risk of cultural homogeneization because of AI because for instance in the music case right imagine uh that most people uh and and there is already evidence for this that most people uh are using these models to generate music for their reals Etc and therefore uh we can expect that uh there will be uh you know this queue will further have a you know uh vicious cycle leading to further marginalization of U uh genres that are under represented there has has been a similar study in Cornell about writing styles and they showed that um you know uh AI generated um sorry generative AI based writing assistants lead to homogenization of writing styles and uh usually these uh writing styles uh are more geared towards certain groups like typically us or Canada uh based writing style and uh the study was between us and uh India two cohorts uh and it showed that how uh you know AI generated writing when used by AI assisted writing when used by Indian uh they really moved away from their natural style much more than what happened in us but in either case there was cultural homogeneization so uh the point is that it is extremely important that uh most uh of the cultures around us are well represented in these llm and the applications that we build so that there is less risk of homogeneization and and this is not something that I'm uh the first person to say uh very interestingly two years ago or or maybe four years ago back in 2020 there was hardly any paper that talked about culture like I already mentioned but uh this graph plots a quarter- wise number of papers that we see getting published uh in the domain of llms and cultures and you see that how it's exponentially growing so there is a big uh group of scientists who are interested in looking at um at the intersection of uh Foundation models and culture more specifically this is for uh llms and culture and the common narrative that emerges in all these cases are that okay llms are biased towards Western culture and they under or misrepresent non-western culture but then uh a natural question is then uh you know okay we want our machines to be culturally aware but uh why uh and when of course I talked about cultural homogeneization but it it depends on lot of factors right how people are using them in what context so it isn't isn't very clear in any of the studies As I Shall argue later with uh you know more data that uh it's not clear why or in what context cultural awareness in machines are very important and that will be the primary uh theme of today's talk so first let me lay out uh the typical research agenda in the papers that uh we see today uh so people ask two kinds of questions uh one is our llms culturally biased and the second is if they are biased how to overcome these biases and to show that llms are culturally biased there are two strategies uh that we observe one is uh what I call culture specific probing so create test sets for a culture so for instance I want to test whether my model understands uh Indonesian culture so I will probably create a data set which asks a lot of questions about Indonesia Indonesia specific questions and I will test different models how well they can answer it and there are already data sets like MML which are more let's say North America specific questions and if you try with them and you can compare you can show okay Indonesian specific questions are not answered that well as us specific questions and you can show that there is a bias or or uh there is this other way of doing uh these studies which is called s social democ graphic probing uh where you ask the model to behave as uh you know personas from different social cultural background like imagine that uh you are a person from such and such country and such and such religion and and how would you solve this task or answer this question and depending on those answers as well you can probably show that there is some sort of biases we we will will look at these examples in some more detail uh but and and once you have shown that there is some bias obviously the way to get rid of cultural biases could be like you try giving prompts which give lot of information about a culture or you can fine-tune with data from that culture right so this has been the strategy so I am going to question uh these strategies like how valid or scalable this strategies are what are we missing out here but before we get in there I will take uh maybe 15 minutes to give you a typical example of this kind of study which I did and this was a work done at Microsoft so in this um particular case um so I was a touring Microsoft uring at that time and I was looking at the ri aspects of um being co-pilot uh that that was just being released around that time and uh one of the thing that bothered me is the policies themselves right so there were a lot of different policies uh like what to be filtered but not to be filtered and uh and it seemed like you know these policies are quite arbitrary uh you know of course these were well thought out and decided by a body or some people but uh they don't apply in all context and everywhere and in different context and different uh applications um it almost seems like these are arbitrarily decided so to test this kind of um I mean ideally we would like the llm itself to somehow or the system itself to have some sort of understanding that based on the application context and the need it could decide what's the ideal h you know policy here should be and we wanted to run a small experiment to see how well models could do what we called moral reasoning and uh the to um create a setup for that and a data set for that we created moral dilemas now moral dilemmas like trolley problems many of you might have heard about already but let me give you a more colorful example of a moral dilemma from the um Hindu epic Ramana so in this epic uh the hero is Rama uh who is seen here uh this person on the right and uh this is the king dharat his father and it's a dharat dilemma dharat went on a Hunting Expedition 10 years ago before this scene is happening and uh he had three queens and he went on The Hunting Expedition with with his second Queen who is standing here uh and uh a tiger attacks dharat and this queen who is very brave she shoots an arrow at the tiger and saves dhat's life so dharat is uh extremely grateful to his Queen and he asks her for anything that she wants because you saved my life so please let me know what you what I can do for you the queen is very smart she said I will ask you when the time comes and dharat says okay granted I promise you to give you whatever you ask for when you ask for it now 10 years later um Rama is now going to be declared as the Crown Prince and dharat had three queens so Rama is the son of the eldest Queen and she is the second Queen so at that time this queen comes uh and tells the sharat that uh you had promised me to give whatever I asked send Rama to Exile for 14 years and make my son bhat as the Crown Prince or the king and um dharat is really Tor right he has a very strange dilemma either he has to break his promise that he made to his queen or he has to send Rama to Exile uh without any fault of Rama so would it be fair on Rama or when whether uh you know uh or or even the people around right the um citizens did not want it either because Rama would be a very ideal King so what should thear do uh now I don't know how many of you already know about ramayana uh if you know or even from the photo you this image you might have guessed that Rama was actually sent to Exile this is typically The Narrative Arc of any epic right that something bad happen with the hero and eventually they emerge uh you know fights all atrocities and challenges and comes back as the hero become the hero so that's what happens here but the question is then why dharat was dharat an unfair King why he chose to keep the promise we'll get to the answer uh in a moment but these kind of dilemas actually happen every day right like for instance who should be the first author of a paper this is such a common thing like even at s i remember uh not between colleagues usually uh these fights would happen between the research fellows who would work with us uh and uh and now I see it with the students who work with me so and and people would come up with different kinds of reasoning for why somebody should be first author and should be not and it's not that some person is right or wrong it it's often it's it's a you know pitching some values against each other so for instance is keeping a promise a greater virtue than being fair to your son and the interesting thing is there is no good right or wrong answer to these questions and often the answers are deeply rooted in culture for instance keeping a promise uh is an extremely important part of certain uh sections of the Indian culture uh and and there are in fact interesting Proverbs uh around it as well which all goes back to these epics for instance uh now uh but but it's not the same everywhere in the world and social scientists have shown this um for instance in this map uh of world value survey there there are many such actually uh attempts to classify by culture and uh and and I would say all of them suffer from problem or the other because you cannot really culture is very elusive to Define right very difficult to Define uh but um each of them take a certain aspect of culture and try to model it in wvs or World value survey uh it's mostly these two Dimensions that they want to capture one is survival versus self-expression which is the x-axis and one is traditional versus secular values and uh countries here like the Nordic countries are high on self expression values so more like individual freedom of and freedom of expression and also secular values and whereas countries here like African and Islamic countries are more about Traditional Values and survival values and of course it's not about which is right which is wrong it's I mean there was for a moment um some social scientists tried to make this claim that this is the axis of you know progress but obviously that's because the you know people who colonized are mostly here and people who got colonized are here so it's a very Colonial View and nobody accepts that view anymore so it's it's very hard to say or or it's a useless pursuit to try to even say which is better than what but it's just a description of how the world is and the point that different parts of the world follows different values now if we want to measure um you know uh how llms do so suppose you ask such a dilemma to an llm uh so this is the story of The Dilemma let's say it's it's a different dilemma but you could imagine dashar story here and you could ask the llm what should dasar do now the problem is this is not a good test because whatever the llm answers it's going to be you can't say right or wrong right because there is no right or wrong here so we changed the test a little bit we added so basically these kind of dilemas exist data sets exist and moral psychologists use it and we added a few more dilemas to it to make it more you know to to spend different values and different cultures uh so so this comes back from 1979 it's called defining issues test but but we changed the test a little bit we added what we call the moral principle here the moral principles States a very in a very specific way uh what is the value that one uh that the person here in question would like to uh uphold and uh this is very uh I mean these principles are defined in such a way that there is a clear resolution for the Dilemma for instance in dashar dilemma imagine I said the policy is a king should always keep his promise that is the most important thing then it's clear what dharat should have done and I could say you know uh the king should not be unfair to any innocent and then probably it will shift the decision in other ways so these policies were designed in such a way that there was a clear resolution so ideally now there is no confusion and if the models are able to do proper reasoning they should be able to uh get to the right answer and what do we observe so we tested uh these five models and they are organized in order of their sizes this was gbt 3.5 turo and gbt 4 and don't look at all these numbers they're confusing just look at the last row uh which is the average and higher the better so 100% means uh everything answered correctly as expected and you would see that gbt 4 is pretty close to 100% which is good it does look like it can uh re quite well gpt3 is close to random even though we had three options uh one is should do it one is shouldn't do it and one is can't say the models never selected can't say it's it's I mean none of the models ever selected can't say so effectively there were two choices so 50% Is Random so gpt3 countries and at all uh not surprising right but what is surprising is as the models grew in size uh they got better so 53 to 70 to 83 to 93 but what happened to Turbo why 68 suddenly there is a drop right and the point is um so then we try to measure uh what we call the bias of a model as follows so suppose a default decision of the model is when you don't give it a policy now if the policy asks the model to decide in the opposite direction so suppose the default was dharat should keep his promise that we observe from a model now we change the policy in such a way that the answer should be he should not keep his promise now is the model able to change flip the decision and uh the less it is able to flip the decision higher is the bias so zero means no bias and uh you know a high number means uh High bias and you see you know gb4 has very low bias which means it can actually flip it its decision based on what the policy is and uh and and it it it slowly grows uh with the model size except for suddenly this Chad GPT GPT 3 it was Chad GPT at that time um which couldn't and this is what we defined as the over alignment problem the problem here is Chad GPT is so overly aligned to certain values that it then becomes very difficult for it to change its stance even if it demands otherwise and if you now ask what this value system is to which it is aligned to no surprises it's typically this kind of value system and in fact um uh there was one dilemma which I did not describe and I don't have time to describe the Dilemma uh that we added rajesh's dilemma uh where actually there was a clear conflict between I mean that the two extremes of the Dilemma uh was one was to decide in favor of uh traditional Society traditional survival and the Other Extreme was secular individual you know self-expression so it was com the two values competing where from these two opposite ends and this was the Dilemma where even GPT 4 could not flip its you know decision so it does seem like uh the models are strongly aligned to certain Valu systems which means if you want to deploy these models in rest of the you know cultures you might be have a problem and you might have to do something extra to make sure that it agrees to the kind of value system that users from these re regions would expect uh now this was the conclusion and of course we did not administer wvs to this is a survey so you can administer it to model we didn't try that but there are other people who tried that and showed exactly what I'm showing in this um you know plot so this is one you know the typical way you show that llms are biased towards certain culture and um I'll quickly take one minute to describe another interesting way where we showed some bias uh in llms but it's not bias in llms it's it's a rather uh bias in the guard rails that we create um so what happens is since there is so much discussion about uh race gender based discriminations uh we go um we do a lot to make sure that llms are not racist llms are not sexist and so on but there are other dimensions of discrimination like cast for instance in South Asia uh which are not so often talked about and therefore the llms don't even have the guard rails from for those uh dimensions and it comes out very uh you know simply in this experiment that we did where we created a scenario where uh it's a interview uh it's a candidate profile discussion not an interview it's a candidate profile discussion so uh hiring meeting rather I would say and uh two colleagues are discussing a candidate and what we Chang are uh the names of the colleagues reflecting certain race certain cast certain gender and name of the person discussed and uh with that we asked the model to generate conversations and other than I mean gb4 again is quite good I would say it did not um derail much but all the smaller models like Lama and mistell easily generated very ctist comments you know comments like they are uncultured and we cannot hire this person because they come from a lower cast and all it just shows that nobody even thought about uh building guard rails uh for these uh other dimensions so other kind of hars that can happen in Societies or cultures which we have not really thought about so this is sort of the state of the art uh now let me go back to the main question s um but but before that is there any question I can stop here for a moment uh no not for now you can continue okay right so um so we saw that we are trying to measure cultural bias using certain strategies so one natural question is then you know we are talking of values yes fine or we are talking of cast but culture is much more complex right so how do we Define culture when we are talking of you know measuring uh cultural bias um and uh is this a robust technique to measure cultural bias I mean for for example the exam you know the two experiments that I mentioned one in detail and one briefly about the cast um are these really robust ways to measure biases in llm uh and and um what what we mean uh by even biases and of course once we can answer these two questions the third question would be um uh when we talk of overcoming cultural biases what do we even expect from the models to do what would be the ideal Behavior so to answer the first question uh we did a study of all the papers published uh till middle of last last year and that they talked of culture studying culture and we tried to see what how they defined culture so this was with respect to the first question what is culture uh and I I would specifically like to call out Jackie um who was a collaborator in this project and um of course my ex- colleague at now in MSR Africa so what we observed is um none of the papers that we studied actually Define culture from an anthropological point of view rather culture is studied through some proxies and I will come to in a moment what do I mean by proxies um the common narrative is of course llms are biased towards some notion of Western culture and the most important Gap that we observe is there is hardly any situated study of actual llm based applications in different cultural settings so uh the the examples that I showed were not really actual applications right and and we did not study users interacting with the system and uh what impact did the bias had on users these are more like model level studies turns out that there is with llms there are at least till uh to best of my knowledge uh no or one or two studies only which looks at you know actual effect on the end user of such biases so this was the biggest Gap now talking about the proxies so whenever we talk of culture it's always defined at a level of a group now this group can be defined by certain sort of um what we call demographic proxies or demographic variables if you will now this could be things like ethnicity which is actually quite a well studied uh Dimension uh in in the literature that we have we surveyed language of course and region in fact region as often it is these days called the geoculture geographical culture is one of the most studied you know Dimensions so and and typically people go with countries right uh and and you can see the problem with that because often uh when you talk of countries like USA or India they're very heterogeneous within themselves so it's very hard to talk about a broad us culture or a broad Indian culture of course you can talk about it but it would be like stereotyping or overgeneralization um then uh the other dimensions which people looked at are gender race religion by the way I'm not talking gender bias here or racism here racist or you know how to put guard rails Against Racism and things like that uh what I'm talking of is um looking at um you know the cultural element in um let's say across races or across genders and study around them uh and religion but there are many things which are not studied or one of the thing which is uh conspicuously Missing From Here is um uh you know Financial level you know social class or financial level which which which or or Urban rural so those kind of cultural because if you think from a user interface point point of view those things might make a big difference especially in certain parts of the world uh the gap between urban and rural users is much higher Gap by by Gap I mean difference so the difference in expectation between urban user and the rural user might be much higher than difference between the two genders for I mean the genders for instance so that that's one thing which is missing the other thing is uh you you know you are talking of a group but what are the things that you are studying about the group because like I said culture is very complex you can't study everything in a single study so some of the things which people seem to have looked at is emotions and values are the most common thing in fact we ourselves looked at values like I said in the example uh and this is probably because values there are data sets like World value survey hofstead value dimension s of different cultures and so on and um there to the uh you know to the mess if I may say because otherwise culture is very fluid um there has been couple of studies about food and drinks social and political relations a bit on technology basic actions and Technology by the way these categories Thompson's 21 dimensions of uh how each culture and there are many such proposals it is one of the proposals there are several other proposals of how culture can be uh what is being studied can be divided so here you see there are a lot of things which are missing for instance um this Thompson's Dimension like you know um social um I mean this uh social relationship there there are lots of details for for instance here which have not been looked at or um um and and 21 there are only five what were the others but yeah there are many more which are missing so that being the state of that let's move to the second question which is are these methods of probing about culture uh reasonable to do uh uh to this end uh we did this study where um we are looking at uh a specific type of probing llms which is called social demographic prompting and we did some experiments I'll come in a moment but I want to call out again here uh Kika and sunna uh were our collaborators and this work was supported by the Microsoft afmr Grant so let's come to uh how social demographic probing works so you start with a prompt like this a person's favorite food is and let's say we plug in a food name so this is what we cultural Cube and then you have how would they solve this and then you have a instruction followed by a question now this question typically comes from a data set so there are many s this part is a biology question and it comes from a data set called mlu a very popular data set to evaluate large language models mostly general knowledge questions about different topics I mean not all general knowledge some are like um math summer biology physics chemistry like covers a whole lot of subjects but supposed to be independent of culture right the answer to this question should not depend on what somebody's favorite food is if it depends then something is wrong with the llm um now what we did was we changed we varied these proxies instead of food sometime we change them with countries names of people kinship Etc so we tried various things and these are highly related to culture so highly culturally sensitive then diseases and hobbies you won't expect culturally sensitive to be culturally sensitive but you know there is a little bit of uh correlation between maybe certain cultures and certain diseases or certain hobbies and certain geographies and so on and lastly these thought were not related to culture at all so I mean not related to geoculture at all like programming languages or uh planet or house number so the question would be if somebody's favorite programming language is python how would they solve it or somebody's house number is 26 how would they solve it so this is just to see how robust llms are to this kind of probing right and in the data sets we tried data sets like mlu where there is these are very factual questions like from science and math so answers should not depend on whatever queue you plug in here no matter is it high medium or low and then for uh there's another data set ethics which is also supposed to be Universal ethics but as I said there is no Universal ethics so you can expect some variation here and finally there are data sets specific to culture so this is called cultural um language inferencing and this is eor a corpus of etiquet so here it should be highly culturally sensitive so what we expect is when we put a low uh low uh sensitive proxy here the answer should not um for so the answer for the low uh sensitive data set should remain fixed no matter what we put here and the answer for a low sensitive propy should not depend on which data set we are proving it with so these are the two hypothesis that we have for a good model and let's see what we see so this is uh llama uh 3 8 billion model and on mmu where it should not the answer should be fixed no matter what food you probe it with right and you see uh I mean this heat map shows the accuracy as you change the food and these Foods were tied to country names I mean we did some rough Association like India was my something like Biryani and I I forgot the details but each each country had a very characteristic food of that country so that's that's how we plotted and if you change it to kinship terms the values again change and you see the colors are also different so the accuracies also change suddenly so what's going on right and this gives you a even clearer picture so this is a four different models uh and uh if you see gp4 because that's the model which agrees with our intuition so let me start with that so MML is this uh purple one and you see variation in MML so what we are plotting here is variance in accuracies when we change the proxy the variation by country by name whatever you change MML answers don't change so that's a good sign that's what we would expect but that's not true for the other other models right you see how much Lama 38 billion changes with country names in a for mlu or even mistal changes then the second thing is for ethics it does change a lot and in fact the ethics data set was supposed to be Universal it just shows that I mean we later looked at the questions individual questions also and it uh we can't call them Universal they are culture specific so here it varies but um uh in interestingly you know it it varies for everything so it varies for Country it varies for disease it varies for hobby so it isn't clear how gb4 comes up with a correlation between a hobby and a ethics question for instance so it's not all good but still better and others you can see it varies with almost everything everywhere quite quite randomly so what it shows is these kind of measures that we have of socio demographic probing to show that a model is biased doesn't tell us much it might be just purious random correlations that we are observing in fact one of the biggest problem uh of measuring bias in llms is because of the maximum likelihood decoding for instance if you ask an llm and this is again a work that I did with Kika and Sena um but yet unpublished if you probe an llm asking it generate about a festival from India it will generate if you ask that question 100 times 90 times it will generate Diwali and 10 times it will generate holy now somebody might look at it and say hey these llms are biased towards um um North Indian festivals or it is biased towards Hindu festivals and there are festivals of other religion that it is not covering but that's not true when you ask llms to generate a festival which is celebrated by keralan Muslims which is a southern Indian state it is able to generate so it's not that there is no knowledge of it it's because it gener it decodes it in a maximum likelihood fashion so it will almost always generate one thing uh which is high probability almost always it will you know kind of generate that so that's why we claim that this kind of uh you know social demographic probing may not be the best way to identify biases in models so then the question is how should we measure cultural awareness and biases in models then and to answer that question I think a more interesting and important fundamental question to ask would be what do we expect from the models and systems right and to answer that question we would need to go one level deeper what real world applications uh demand cultural awareness Because unless we have an application we can't answer what we expect from systems and models and unless we answer that we don't know how to measure it right so that should be I think the strategy and in the last three minutes I will quickly talk about um something I mean we are in a very early stage of answering those so I have more questions than answers but I'll just tell you how we are thinking about it so to start with we thought we will build an application which requires some sort of cultural awareness so the application we built is we call it culturally yours it's a cross-cultural uh reading assistant so imagine you are reading a good reads review of a book by an Ethiopian author and the review is also written by somebody from Ethiopia and now you are not from Ethiopia let's say you are from the US you have some knowledge of Ethiopian culture but not a lot then there might be things in the review that you really don't understand so what culturally yours try to do is knowing your background suppose it knows that you are from us and you are such and such year old it will ask you one or two basic questions about your demographic variables or proes as we introduced it earlier and based on that it will try to automatically highlight things which you may not understand and once you um you know hover your mouse over uh what it highlighted it will try to explain you what that is note that it might be wrong right in the Assumption for instance it talked about the article talked about inera bread and you are a foodie you know Ethiopian food uh you may not know Ethiopian monuments but you know Ethiopian foods so uh you know you might say I know this once you give that feedback the model will try to personalize it will uh at the back end try to build uh a more precise model of your culture uh and and which is which goes much more Beyond than your country and your age and your education level for instance right your personal interest and a lot of things shapes your knowledge and understanding now before I show you uh some interesting uh yeah before I show you some interesting you know uh theoretical things I want to mention one user study that we did with good reads review from people from three um countries USA India and uh I think Mexico uh evaluated um you know marked books from book reviews written by USA India and UAE as what was difficult to understand for them and then we annotated them by linguistic difficulty cultural difficulty and so on and 50% of the difficult to understand spans more than 50% come actually because of cultural uh you know mismatch so I mean or lack of cultural knowledge uh so you already see that there is u a need for this kind of a tool we haven't worked much on the back end I must say like we are just using simple prompting approach with uh models like gb4 but again gb4 being very good uh we do get reasonable results uh but but what I want to talk about about uh uh here is what I just told you is essentially a system which uses culture as a prior for personalization the this reading assistant is useful only when it is very very personalized right to start with it is not very useful probably but if it starts with culture as a prior it still gets certain things right and then it it quickly can you know personalize and there is a theory uh it's proposed by hofstead uh back in 94 which is that you know so there is a universal human nature uh and culture is something specific to a group which we belong to and we each have individual interest and personality on top of it and essentially we are just right you know using this pyramid when I'm talking of culture as a prior for personalization and let me end with this uh you know uh thought that uh since culture is always evolving and since um culture has a long tale right you could talk of culture of amirati I mean all women but you could specify you know it I mean drill it down to Middle Eastern women to amirati women and amirati women will share something with amirati men but will also not share something with men but rather share with other Middle Eastern women and so on so it's a very complex intersection of lot lot of different uh demographic attributes that we all live in right live with so then how do you expect a model to learn or know all of this and the question we asked was how do humans do it right humans also should ideally not understand and know all these complex things and it so turns out that there is this theory of metacultural awareness and it says uh to be a good metac culturally aware so metacultural awareness is awareness not about a culture but General awareness about how cultures behave uh or cultures kind of interact and um there are three fundamental principles of metacultural awareness um and and a individual who pauses all these three are able to navigate in any culture needless I mean irrespective of uh how much aware they are to that culture so first is variational awareness which is given one particular situation what could be all the different possibilities and what is the prior probability uh that you assign and how much is your confidence on that so how good is that to an individual is the variational awareness for instance if I ask you which side of the road people drive in Kenya you may not know the answer but you know that the space of possibili left and right you won't say it could be top bottom down middle no it will be left or right one of those two so that's your variational awareness and you might go one step further okay I don't know the answer but Kenya was a British colony British drives on right so sorry left so maybe it is left so that's one level higher so you have a better prior probability but you also have an uncertainty so quantifying those things are very important and we should measure models on these kind of things rather than just what it is generating you know uh how to do it I don't know yet we don't have a clear answer yet second is then once we don't know something uh or we have uh confusion there should be a way to state that so in uh human human communication if I'm in in a new culture and I don't understand really something I might say that sorry I don't understand this so uh I'm I'm not rude and uh simply I just don't know this and I'm curious to know this so This is called explication and third is negotiation which is either through observation indirect observation or through direct questioning or some sort of probing how can you identify the or understand the cultural elements from the foreign culture with in a very sample efficient way and uh once you have all these three qualities you are set to navigate through any culture so ideally our models and systems should have these qualities and the argument is the models should have variational awarness where the whereas the system level things should be more around explication and negotiation our and our LMS are very bad at explication right they never say what they know what they don't know they can't give the uncertainty quantification well and and they hardly ask questions um so this is where uh um we are right now and this is what I'm excited about and I'll end with this note that um you know it's a different thing to study a culture from outside and a very different thing to live within a culture and uh have a lived experience of a culture and uh unfortunately a lot of the data that the models are trained with are written by people from certain demography and certain cultures and therefore uh even if we have data and uh description and knowledge of other cultures those are typically thin descriptions written by others and not people from their own lived in experiences that's that creates an inherent uh problem and uh I'll stop here I'll not uh go through these uh two questions you can read it for yourself I just put put put out some open questions here when I'm I'm happy to take uh questions thank you

Original Description

Language, culture, and technology are deeply intertwined; any technology that aims to process or generate language must have a nuanced understanding of culture. Yet, culture is elusive, often defying concrete definition, which poses a profound challenge to NLP and AI research. While the NLP community has acknowledged this fact and taken initial steps —developing culture-specific benchmarks and identifying biases in large language models (LLMs)—there remains a gap in addressing why cultural understanding matters for machines, and the real-world impact on users of culturally (un)aware machines. In this talk, I will ask and attempt to answer the following three questions: (1) What can we do better with AI systems that are culturally aware? (2) Why is modeling culture so complex, and what can we learn from psychology and anthropology? (3) How can LLMs, in turn, be leveraged as tools for studying and enhancing our understanding of cultures?
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Microsoft Research · Microsoft Research · 0 of 60

← Previous Next →
1 Frontiers in ML: Learning from Limited Labeled Data: Challenges and Opportunities for NLP
Frontiers in ML: Learning from Limited Labeled Data: Challenges and Opportunities for NLP
Microsoft Research
2 Frontiers in Machine Learning: Climate Impact of Machine Learning
Frontiers in Machine Learning: Climate Impact of Machine Learning
Microsoft Research
3 Frontiers in Machine Learning: Security and Machine Learning
Frontiers in Machine Learning: Security and Machine Learning
Microsoft Research
4 Hope Speech and Help Speech: Surfacing Positivity Amidst Hate
Hope Speech and Help Speech: Surfacing Positivity Amidst Hate
Microsoft Research
5 Early Indicators of the Effect of the Global Shift to Remote Work on People with Disabilities
Early Indicators of the Effect of the Global Shift to Remote Work on People with Disabilities
Microsoft Research
6 Remote Work and Well-Being
Remote Work and Well-Being
Microsoft Research
7 Challenges and Gratitude of Software Developers During COVID-19 Working From Home
Challenges and Gratitude of Software Developers During COVID-19 Working From Home
Microsoft Research
8 Towards a Practical Virtual Office for Mobile Knowledge Workers
Towards a Practical Virtual Office for Mobile Knowledge Workers
Microsoft Research
9 Impact of COVID-19 crisis on the future of work in India
Impact of COVID-19 crisis on the future of work in India
Microsoft Research
10 Empowering and Supporting Remote Software Development Team Members through a Culture of Allyship
Empowering and Supporting Remote Software Development Team Members through a Culture of Allyship
Microsoft Research
11 How Work From Home Affects Collaboration: Information Workers in a Natural Experiment During COVID19
How Work From Home Affects Collaboration: Information Workers in a Natural Experiment During COVID19
Microsoft Research
12 Phong Surface: Efficient 3D Model Fitting using Lifted Optimization
Phong Surface: Efficient 3D Model Fitting using Lifted Optimization
Microsoft Research
13 Managing Tasks Across the Work-Life Boundary: Opportunities, Challenges, and Directions
Managing Tasks Across the Work-Life Boundary: Opportunities, Challenges, and Directions
Microsoft Research
14 Microsoft Urban Futures Summer Workshop | Data Driven Urban Transformation [Day 1]
Microsoft Urban Futures Summer Workshop | Data Driven Urban Transformation [Day 1]
Microsoft Research
15 Microsoft Urban Futures Summer Workshop | Sensors and Data [Day 2]
Microsoft Urban Futures Summer Workshop | Sensors and Data [Day 2]
Microsoft Research
16 Microsoft Urban Futures Summer Workshop | Policy and Social Impact [Day 3]
Microsoft Urban Futures Summer Workshop | Policy and Social Impact [Day 3]
Microsoft Research
17 Directions in ML: Algorithmic foundations of neural architecture search
Directions in ML: Algorithmic foundations of neural architecture search
Microsoft Research
18 MineRL Competition 2020
MineRL Competition 2020
Microsoft Research
19 Can we make better software by using ML and AI techniques? With Chandra Maddila and Chetan Bansal
Can we make better software by using ML and AI techniques? With Chandra Maddila and Chetan Bansal
Microsoft Research
20 From Paper to Product
From Paper to Product
Microsoft Research
21 SkinnerDB: Regret Bounded Query Evaluation using RL
SkinnerDB: Regret Bounded Query Evaluation using RL
Microsoft Research
22 From SqueezeNet to SqueezeBERT: Developing Efficient Deep Neural Networks
From SqueezeNet to SqueezeBERT: Developing Efficient Deep Neural Networks
Microsoft Research
23 Programming with Proofs for High-assurance Software
Programming with Proofs for High-assurance Software
Microsoft Research
24 Platform for Situated Intelligence Overview
Platform for Situated Intelligence Overview
Microsoft Research
25 Directional Sources & Listeners in Interactive Sound Propagation using Reciprocal Wave Field Coding
Directional Sources & Listeners in Interactive Sound Propagation using Reciprocal Wave Field Coding
Microsoft Research
26 Galactic Bell Star Music Demo
Galactic Bell Star Music Demo
Microsoft Research
27 Importing Animations in Microsoft Expressive Pixels (9 of 9)
Importing Animations in Microsoft Expressive Pixels (9 of 9)
Microsoft Research
28 Welcome to Microsoft Expressive Pixels (1 of 9)
Welcome to Microsoft Expressive Pixels (1 of 9)
Microsoft Research
29 Getting Started with Microsoft Expressive Pixels (2 of 9)
Getting Started with Microsoft Expressive Pixels (2 of 9)
Microsoft Research
30 Creating an Image in Microsoft Expressive Pixels (3 of 9)
Creating an Image in Microsoft Expressive Pixels (3 of 9)
Microsoft Research
31 Creating Animations in Microsoft Expressive Pixels (4 of 9)
Creating Animations in Microsoft Expressive Pixels (4 of 9)
Microsoft Research
32 Managing Animation Galleries in Microsoft Expressive Pixels (5 of 9)
Managing Animation Galleries in Microsoft Expressive Pixels (5 of 9)
Microsoft Research
33 Creating Fragments in Microsoft Expressive Pixels (6 of 9)
Creating Fragments in Microsoft Expressive Pixels (6 of 9)
Microsoft Research
34 Using Layers in Microsoft Expressive Pixels (7 of 9)
Using Layers in Microsoft Expressive Pixels (7 of 9)
Microsoft Research
35 Exporting Animations with Microsoft Expressive Pixels (8 of 9)
Exporting Animations with Microsoft Expressive Pixels (8 of 9)
Microsoft Research
36 What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 2/2)
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 2/2)
Microsoft Research
37 What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 1/2)
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 1/2)
Microsoft Research
38 Planeverb: Interactive sound propagation for dynamic scenes using 2D wave simulation
Planeverb: Interactive sound propagation for dynamic scenes using 2D wave simulation
Microsoft Research
39 Making cryptography accessible, efficient, and scalable with Dr. Divya Gupta and Dr. Rahul Sharma
Making cryptography accessible, efficient, and scalable with Dr. Divya Gupta and Dr. Rahul Sharma
Microsoft Research
40 Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 Talk)
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 Talk)
Microsoft Research
41 Optics for the cloud – Light at the end of the tunnel? (SIGCOMM 2020 Workshop)
Optics for the cloud – Light at the end of the tunnel? (SIGCOMM 2020 Workshop)
Microsoft Research
42 Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 short talk)
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 short talk)
Microsoft Research
43 Sirius: A Flat Datacenter Network with Nanosecond Optical Switching (SIGCOMM 2020 short talk)
Sirius: A Flat Datacenter Network with Nanosecond Optical Switching (SIGCOMM 2020 short talk)
Microsoft Research
44 Novel Image Captioning
Novel Image Captioning
Microsoft Research
45 Forest Sound Scene Simulation and Bird Localization with Distributed Microphone Arrays
Forest Sound Scene Simulation and Bird Localization with Distributed Microphone Arrays
Microsoft Research
46 Decoding Music Attention from “EEG headphones”: a User-friendly Auditory Brain-computer Interface
Decoding Music Attention from “EEG headphones”: a User-friendly Auditory Brain-computer Interface
Microsoft Research
47 How does holographic storage work?
How does holographic storage work?
Microsoft Research
48 The physics of hologram formation in iron doped lithium niobate
The physics of hologram formation in iron doped lithium niobate
Microsoft Research
49 Introduction to coax: A Modular RL Package
Introduction to coax: A Modular RL Package
Microsoft Research
50 Directions in ML: "Neural architecture search: Coming of age"
Directions in ML: "Neural architecture search: Coming of age"
Microsoft Research
51 Microsoft Research AI Breakthroughs 2020: 20 minute research talks + Q&A panel
Microsoft Research AI Breakthroughs 2020: 20 minute research talks + Q&A panel
Microsoft Research
52 Fireside Chat with Johannes Gehrke during Microsoft Research AI Breakthroughs 2020
Fireside Chat with Johannes Gehrke during Microsoft Research AI Breakthroughs 2020
Microsoft Research
53 Fireside Chat with Susan Dumais during Microsoft Research AI Breakthroughs 2020
Fireside Chat with Susan Dumais during Microsoft Research AI Breakthroughs 2020
Microsoft Research
54 Microsoft Research AI Breakthroughs 2020: 20 minute research talks, Q&A panel, and event wrap-up
Microsoft Research AI Breakthroughs 2020: 20 minute research talks, Q&A panel, and event wrap-up
Microsoft Research
55 Clinical Research with FHIR
Clinical Research with FHIR
Microsoft Research
56 Soundscape Street Preview
Soundscape Street Preview
Microsoft Research
57 Tilt-Responsive Techniques for Digital Drawing Boards
Tilt-Responsive Techniques for Digital Drawing Boards
Microsoft Research
58 SurfaceFleet: Exploring Distributed Interactions Unbounded from Device, Application, User, and Time
SurfaceFleet: Exploring Distributed Interactions Unbounded from Device, Application, User, and Time
Microsoft Research
59 Haptic PIVOT: On-Demand Handhelds in VR
Haptic PIVOT: On-Demand Handhelds in VR
Microsoft Research
60 SurfaceFleet Supplemental Video Demonstration (UIST 2020)
SurfaceFleet Supplemental Video Demonstration (UIST 2020)
Microsoft Research

This video teaches the importance of cultural awareness in language models and explores the challenges and limitations of current approaches, highlighting the need for more nuanced understanding of culture in LLMs, with tools such as GBT 4 and GB4 being utilized. The video discusses the concept of culturally aware machines and the need for LLMs to be able to reason morally in context. The speaker presents a moral dilemma from the Hindu epic Ramayana and notes that the answer to the dilemma is de

Key Takeaways
  1. Create test sets for a culture to test different models
  2. Ask the model to behave as personas from different social and cultural backgrounds
  3. Give prompts with a lot of information about a culture to try to get rid of cultural biases
  4. Fine-tune with data from a culture to try to get rid of cultural biases
  5. Use social demographic probing to test the robustness of LLMs to culturally sensitive prompts
💡 Cultural awareness in machines is not clear why or in what context it's important, and there are two strategies to show LLMs are culturally biased: culture-specific probing and social demographic probing.

Related Reads

Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →