Why Am I Seeing This?

Data Skeptic · Beginner ·⚡ Algorithms & Data Structures ·10mo ago

Key Takeaways

The video 'Why Am I Seeing This?' by Data Skeptic explores the challenges of studying social media recommender systems and introduces the 'recommender neutral user model' for inferring the influence of recommenders on user interactions, using tools like DBLP dataset and computational models.

Full Transcript

[Music] Welcome to Data Skeptic, a podcast exploring the methods, use cases, and consequences of recommener systems. Well, welcome to another installment of Data Skeptic Recommener Systems. Today, we're digging into one of the trickiest problems in social media research. You can see what people click, but you can't reliably know what they've been shown. That missing piece of information shapes what we know about user preferences, echo chambers, and polarization. Most of the best known recommener systems are in industry. They're proprietary systems. And while those teams might sometimes publish, there's no guarantee we know exactly what's going on in their production systems. Our guests today are tackling that blind spot headon. Joining us are Sabrina, Dmitri, and Gregor, whose work explores how to infer the influence of opaque recommener systems, even when the platforms don't share their exposure data. You'll hear about their recommendder neutral user model. It's a graphneutral approach that learns user behavior without baking in any single platform objective. Their methods allow them to simulate how different recommendation strategies would have steered users, even without the privileged access you would like to have to all the logs of the platform. We also zoom in on the stakes. Why optimizing for engagement can backfire on users. Essentially, that's Goodart's law that we've talked about here before on the show, how feeds can accelerate polarization, and what practical fixes might look like. So, if you're building, auditing, or simply living with recommener systems, this conversation offers tools and a mindset for seeing beyond that click. Let's jump right in. >> I'm Sabrina Guidi. I'm a master student in computer science at the University of Milano Bikoka. >> I'm Gregor. I'm a PhD student at the University of Ringsburg which is in Germany. >> And I'm Din. I am the director of the BCON lab in University of Milano Bikoka in Milan, Italy. >> And how did your collaboration between the three of you begin? Actually, I was a supervisor for Sabrina undergrad projects and then she's doing also her master work with me and Gregor I think we met in a project about studying the impact of social media on teenagers and Gregor did a lot of work on that. >> No. Yeah. And so I think Udo who is my supervisor he spent some time together with you in England right and that's where you met and then there was this contact between the two of you and that's where I came in. >> Yeah we were working together in uh UK in Essex for some time. >> Well we're going through on this podcast a series exploring lots of interesting research related to recommener systems. Is recommener systems a focus for any of you guys or is this sort of fits in with a bigger scope of research? So actually for me is a very important topic. We are doing several project on recommener systems because I think they actually can have a strong impact on user experience online and uh in turn this affect what we know about the the world and the decision we take. So I spent some time working on this. We got some grants. There was this grant for Volkswagen foundation um was called it coverage and we were trying to study actually how different commanders would affect the users behaviors the formation of echo chambers built bubbles and so on and how this will led to polarized opinions and um but actually it's very challenging because there are not many data online so yeah the the work is going on I think will last for quite a bit. >> Are there any canonical data sets you can look at or do you have to create your own data? >> So actually there aren't a lot of data set about this. For now there aren't any data set that both record the user interaction history in a social media and which recommener system affected them I mean influenced their decision. And because of this we worked on sitation data set in particular DBLP and because its structure is very similar to general purpose social media and also the data is fully available. So we used this type of data in our research. >> So actually I I know there are a few data sets but they are mostly not for social media. there are data on recommenders interaction with users on I don't know for example Amazon uh with reviews and things like that but they are missing uh important aspect for us like social connections that is important the the way you get an information if it's coming from an advert or is coming from your friends really change the way you will interact or will take into account this information. So most of the data sets that are around are not really suitable for this and and this is why we are actually trying now to collect the data set. We are creating a platform to collect data set for example from the fediverse from mastered on uh platforms so that we will have anonymized data but we can study the impact of the recommener used on some of these platform because now on fiver some of the platform actually use recommenders to us in classical social media and also we are trying to collect similar data from other sources like classical social media platform. and asking users to share data after a running mid session. So yeah, it's not easy to get data to tackle this problem but as it can affect a lot of social aspects. I I I was in UK like was saying we had the Brexit and this affected us quite a bit and there is all the study analytica. So we want to see what was the role of the recommenders in this kind of matters because recommenders is something we can change. If there is actually a strong impact of an algorithm, the algorithm is something you can change. Changing people, educating them is something we are trying to do but it's much harder than changing an algorithm. I think that's why we are trying to really push on this problem. >> Yeah. One of the challenges is that all the major social media platforms are private companies and therefore they're closed. You know, their data is their own. Maybe some of them have, you know, some research affiliations or something but not generally. Before we get into the main paper, why am I seeing this? Can you say a little bit about that project? Uh, okay. Yes. This was a international project where we participated together with the University of Regensburg, HRW that is another university in Germany and a university of Pompeo Fabra in Barcelona. There we were actually studying how to integrate educational activities into social media and also educating teenagers to make them understand better the impact of recommener system on their choices or social media on their choices. You know there are many issues that are related to social media but also some advantages. For example, during COVID, social media helped people being together, being connected and this helped people. But on the other side, there are various issues like proliferation of conspiracy theories, fake news. Also there is a lot of bullying or there is an impact also on self-image because people see online uh content that seems they can relate to like these strongly polished images of influencers and this affect their self-image. they feel not at the same level and so there there is a lot of work to make people more resilient to the content that is shared on social media and also teach them that when they share something on social media they actually are going to affect other people. So sometime we forget that social media is a not only a passive process actually we affect other people. So some of us are doing the bullying. It's not only we are not only receiving it and people have to learn to not to do that. They also have to learn that when they are liking and sharing things it will affect other people and so that was the the objective of the work and again because most of the data on social media platform are from private companies and for them is very important to keep those data private. It is hard to study what is happening. So in many cases what we have to do is to use computational models of social media dynamics of uh how people take decision how people are affected and try to get data to validate this model. So it's not really not really simple but uh I think it's important so we are trying to go in that direction. >> Well there's a term I was hoping I could ask you to define from the title. The title being why am I seeing this towards recognizing social media recommener systems with missing recommendations. What is a missing recommendation? >> Because when we use only observed data, we do not know which recommener system the platform actually use. So we are trying to guess which social media is the most likely to be used in this context. >> Okay, maybe I can give you an example I use often. The problem with the current most of the current data set on social media is that in many cases they only present the actions that the user took. So for example, let's say that Instagram is showing me 1,000 picture of cats and only one picture of dog. So if I choose cats, then the only thing that will end in most of the data set is my choice. there is no information about this imbalance between the content I'm shown and the actions I do. So when people analyze this data and they come up with things like ah users prefer cats to our dogs online if the data are skewed like this if people are showed 1,000 cats and only one dog is not really the user choice it is the consequence of the algorithm. So when when we hear all these news about people behavior online this actually should take into account what the algorithm are doing because they are presenting the users as skewed virtual reality because they are following some strategy that we don't know and then this is what leads users behavior. No. So what we are missing are this recommendation. We don't know usually what has been recommended. We don't know this imbalance between the 10,00 cats and the dog one only dog feed that was presented. Okay. So it is quite challenging to get something out in this condition. >> So there are many different algorithms a company might use and uh I guess some companies will publish things about their methods but they're not required to. They could have a private algorithm they're applying. How much insight do you have into the algorithms companies use? So actually for example for YouTube for some time there was quite uh a clear understanding of the architecture of the recommenders. Some time ago we also had some kind of leakage of data from Twitter. I don't know if you remember that there was a a story about the code where there was if the owner of Twitter is posting this everybody will receive it or something like that in in the algorithm but actually in many cases even if you have the sodo code or the code if you don't have like the objective function and if you don't have the data that they're using to train it it is very difficult to predict what's going on okay because in many cases most of these model are big neural networks and if you don't know the reward functions or what is being improved and what is being punished what is the objective of the algorithm it is hard to know what is going on a lot of work that you can find in academia about recommenders is not really about trying to predict their behavior make it accountable or so on but is to make it efficient the recommener system of YouTube or Netflix they have to answer millions and millions of users uh in a very fast manner over thousand of content or billions of contents. So their challenge and what they publish about is let's say usually performanceoriented but there is not much about the actual impact on the user and they often consider the reward function that they are using the objective as not something they publish. Okay, it's like a function that we put there and that uh and the system will perform well on it. But we don't know for example there is a advertising campaign or this kind of things. It's not really easy for us to know what what has been put there. I think now there there have been some regulations about these kind of things. you know that uh the European community has been or commission I don't know exactly who does it uh has been finding some of the big companies about hidden uh advertisement in this kind of content yeah are not the details that we really can see in uh the algorithms that they publish >> maybe what what I would like to add here is um even if they publish a paper um it there's quite the high chance that the algorithms they actually run on their platform have already changed. So I think that they are making a lot of changes to these algorithms and for example they're also um constantly running AP studies to check whether new components result in higher um interaction on these platforms and stuff like that. So yeah, probably by the date you read a paper that was published by, I don't know, Tik Tok or whatever, it will already be outdated what they published about the algorithms. That makes sense. And I have some personal experiences that lead me in that direction too that recommener systems I've either worked on or worked at a company that had it. It's constant experimentation. What challenges does that pose for you as a researcher? How can you get a head around what people are doing when they're always changing? So actually this is a big challenge and obviously Sabrina has been working with us since not too long and we are actually trying to simulate this kind of variations with our data. What we try to do is to create synthetic data that reproduce what will happen on a social media changing the different recommener system. So actually we didn't do experiment on this but this is something we planned. We will try to see what happens. How much an algorithm that is trying to track the recommendation algorithms will lag behind the actual recommener system is uh that is being adopted because uh you can imagine that anyway what is going to happen is that you estimate what is the right algorithm over a set of observation over some time. No. So uh before the any algorithm decide that there is a there was a switch of the recommendation algorithm it will take some time. So this is another another challenge that we study using synthetic data as we have done for the moment for with static algorithms >> and can you share a few details on that synthetic data generation process? How do you develop a rich enough synthetic data set to be useful for research purposes? >> To do this, we um actually created a recommener neutral user model. So the goal of this is to model user behavior in a way that minimize the influence of any specific recommener system. And we want to learn baseline behavior that isn't biased by any recommener system. And to build this we used like an insight infosphere. So it's like a structure that approximate what a user was likely to to receive certain actual recommendation and directing back from his actual future interaction and identifying the parts that connect this content to him. And then we combine the information with his history to create a richerformational context. And then we train a graph neural network to predict his actual future action. And we can use this model and having as an input different type of recommener. We can recreate the possible user behavior just changing the input. >> Makes sense. Yeah. Yeah. And then I think you've got a nice scalable process to work with in that sense. Delete me makes it easy, quick, and safe to remove your personal data online at a time when surveillance and data breaches are common enough to make everyone vulnerable. Did you know your data is a commodity? Right now, countless data brokers are collecting your name, contact information, and even your home address to sell online. This information in the wrong hands can lead to serious problems. But removing your data will help you protect yourself from these threats. Have you ever been the victim of identity theft? If you haven't, you probably know someone who has. I started using Delete Me after a close friend was harassed by someone who purchased their information online. Delete Me has given me peace of mind knowing my information is being removed from hundreds of data broker websites. Take control of your data and keep your private life private by signing up for Delete Me now at a special discount for our listeners. The only way to get that 20% off is to text data to 64,000. That's data data to 6400. Message and data rates may apply. Imagine a world where AI agents collaborate as efficiently as humans do. That world is here. The agency is an open-source collective breaking down barriers between isolated AI systems. We're building the internet of agents, a universal collaboration layer where intelligent systems seamlessly discover, communicate, and solve complex problems together. For developers, this means freedom from proprietary limitations and access to standardized tools that work across any framework. With backing from industry leaders like Crew AI, Lang Chain, Lambda Index, and Cisco, we're creating the foundation for truly interoperable multi- aent systems. The agency isn't just releasing ideas. We're shipping code, specifications, and services ready for implementation today. Build with other engineers who care about highquality multi-agent software. Visit agency.org org and add your support. That's agency. That's a g ntcy.org. Could you talk a little bit about the simulation of the user behavior? I think I have a good sense of how you can come up with the data, but how do you model the user and all the preferences they have? >> I mean, for noodies, we we use the recommended neutral user model. So I mean it encodes what the user preference during the learning process. So it's actually a model that learns to behave like the user under a different recommener system assumption. >> Okay. So practically what Sabrina did was to learn a user model directly straightly straight from the data. So ply there is a graph neural network that learns to predict what will be the next preference of the next action of the user directly from from data from what will be the actual real future actions and we uh feed this neural network with an input similar to what the user will actually have and this is like his previous experience. So his previous connections what he exchanged with other people that really fit well with the graph neural network structure and also with the recommendations that are provided to it. So in this way we learn a model of what the user actually does. The challenge as I said before is that we during training this model don't know what are the real recommendation that the user saw. So when we create a user model, the challenge is how you know to what the user reacted if we don't know what was suggested by the recommener. So uh Sabrina created what he call insight infosphere. So is a simulated the recommended that should be quite effective on this and it practically takes what actually was executed in future by the user and adds uh similar noisy parts that are similar content that probably the user was exposed to by the recommener. Okay. operating imagine that from what you did till now you have to take three steps to get to the next like you make because it's a friend of a friend that shares that. So the recommener to find that data should do a navigation in the connection graph that is of three steps. So we do similar things. We replicate these exploration paths from the user history from the user connections for a few steps and add these uh similar parts to the one that actually are true. So the context that we provide to the graph neural network to to the user model, the information that the user model has is richer than the actual final information contains information that ma it must learn to discard and it must learn only to select the things that actually were really chosen by the user and learn to discard things that are similar but the user didn't consider. In this way, we can balance between content that the user is likely to be exposed by a recommener and content that actually will interact with. >> And do you have a sense of what the average user thinks about recommendations? Are they even aware that they're taking place? Do they appreciate them or dislike them? What sort of a general sentiment? >> So, this is for example a previous study we did. Uh I think also Gregor is involved in this and this was done with some schools I think in Italy but maybe in Spain too. I don't know if in Germany. So actually we had some surveys for high school students. At the same time we had them do a game where we showed them for example two conditions. In one condition they had content that was provided by a fail recommener. recommener that was actually reproducing information that reflected the average distribution of this information and another recommener that will show them information that was more similar to their opinion and then they had to choose make a decision on the base of type of the recommendation they got and then we showed them that if they use this kind of biased information they will make more errors. Okay. And the interesting thing was yes, they learned about recommenders. They uh knew a bit about it, but the annoying thing is that my my objective with that experiment was to lead users to be more uh aware of the negative impact of recommenders and ask for regulation, ask for recommenders that may be more fair and realistic with their content they present. But actually most of the user kept wanting recommendations that are closer to what they like instead of fair recommendation. So this I think is a surprising feeling that even if you show that with a biased recommendation you make more errors your estimations are bad your decisions are bad people still prefer apparently having this biased recommendation. Okay, this was I think a small set was published I think with the title is something like eco chambers affect you too. It is a small data set I think some hundred people but it shows that people really like to avoid to have to look and read the things that they are not totally interested and and this is what make us more vulnerable to social media recommendations. Well, I think most people who work on recommener systems, and you could tell me if you disagree, but I think most of them are well-intentioned. Their goal is to surface better recommendations for a user, which in theory the user should like. If uh they like them, they'll come back more and click more and things like that. If they don't, maybe they'll go elsewhere. Where does a negative outcome arise from? How do those errors get made? >> This is not always errors. I think in many cases an issue can be the type of reward function you are trying to follow. So nowadays most of the recommenders mix neural models that are kind of blackbox. So features that we don't understand of the content with features that are a bit more clear and easy to drive. And this is how we create more explicit user models. But in the recommener for social media often are mostly based on uh engagement and for example one interesting thing is that is what is called that inflames online. You know, when there is a big fight online and people are really interacting a lot and getting upset a lot and keep fighting and uh for the recommener system, this is not really different from people that are liking a new song because they are going there commenting and interacting and spending time there. So, this brings an interaction that is not really positive. People are getting more upset, more biased, more polarized. But for the recommended is not really easy to understand and probably if you are just trying to maximize engagement. So how much pe time people spend on the social because let's say this is what you think is good for them is they are enjoying it. So you are trying to maximize the time you spend there. So this brings these negative effects that uh can be can can happen. So yeah, I understand that the who develops the algorithms for sure 90% of the cases is well intentionate but these small dynamics can already bring a lot of issues and maybe also connected to what Dimidri just said because I think that is quite interesting when you like optimize your recommendation algorithms on the interactions or like the the time that users spend on social media. it might not be the optimum for making them happy because I I recently read this paper which was called what measures in a metric and they mentioned good hearts law which essentially says when a measure becomes a target it becomes a bad measure. So if you measure like the interactions on social media and the time that users spend there but actually it does not make them happy anymore it's also maybe the wrong thing to optimize on and I think that was also one of the examples that they talked about. So I don't know what pl platform it was maybe also one of the big platforms they optimized the recommendation algorithms and it turned out that it had negative effects on the user satisfaction but increase the time that users spent on these platforms. >> Okay now the two words on that. So another project we have on social media is about what is called social media addiction but actually the technical word is different because it's not yet recognizes addiction is social media problematic usage and actually yes in the end it's just people spending too much time on social media and uh we have we are working on recommenders that try to balance engagement so people should use the social media platform in a sustainable manner both for the platform and for the users and we are not the only one. I remember there is a paper from Meta I think one of the author is I think Loren Soladzerik and they actually try to model ways to avoid the people spend too much time on social media because the the issue in that case if they spend too much time they may drop out because usually when you don't manage well something the solution is drastic you just cut it and this for a platform is not good too so they want to lead people to have a balance at the usage and this is also what we are trying to do in another project. >> Well, if I were working on a recommener system project and I got the result you mentioned earlier, I think it was Gregor where the user satisfaction went down but the time on site went up. This feels like a big conundrum because I'm worried the finance people are going to be happy with more time on site, more ads to show or something like this. But at face value, if users are less satisfied, that doesn't seem good. What advice would you give for someone who gets that result? How do they approach it? >> This is a very hard question I think and the fact that already meta is trying to find ways to balance this is trying to tackle this over usage problem means that they are some way aware of it and yeah probably you have to find a solution where you maintain the platform sustainable. We are trying to do something similar and uh the paper is going to be submitted in a couple of days so I cannot say too much about it but yeah trying to find ways to keep people engaged with the social media platform for uh enough time that you can still get the money to make it run and they can keep going on with their life is I think something important for for them. Uh yeah, there there is interest on that. I think now >> do you think we can solve it with just coming up with the best possible reward function? >> Uh reward function uh it's not the only thing >> or maybe there's more to it. Yeah, >> obviously we we are also working on this educational aspect where we make people aware of uh the complexity and the effect of social media. But I think improving the algorithm algorithmic side is important. they the algorithm are very fast very they have a lot of information and uh and people we we are weak I don't know there there are some papers in psychology that suggests that thinking too much that with training and with what you call call uh willpower you can really fight back with addictions and other issues like that is a bit of wishful thinking people get tired and often after an effort to get out of an issue. I'm talking of other stronger forms of addiction often they have relapse and they go back to addiction. So uh we need to make the environment safer for people so that uh it's not so easy to fail to fall in this kind of uh traps let's say so yeah I think the algorithm should help the algorithm also the interface itself you don't need a really I don't know a general artificial general intelligent to help you you don't need hair as in the movie that helps you but uh you may just have a a notice look you have been spending he has so much time. Are you sure you're going to sleep enough? I think there are this kind of nudging things that can help people realizing that they have to behave in a different way. So it is other than the reward, it is actually the user interface that must be helpful and lead to more healthy usage. >> What role does uh transparency play in this process or what role would you suggest it have? The problem is in many cases like you were saying the finance people or the interest will be a bit different from uh the user interests at least in the short term and uh we don't know what's going on till uh because we are not don't have any access to these recommenders. We don't know if there is any campaign that is being played. This is particularly important for example in politics. know there is a lot there was a lot of discussion about the impact of social media on elections. So being aware of how information is being shared of how you are exposed to other people opinion. Imagine this. It's not true that being exposed always to people with different opinion from yours is good and it's not true also the opposite that you are exposed to opinion similar to yours because that is like yesmen. No, if you meet Yesmen, you believe in more even more in what you already believe. But uh if you are exposed to people with a very different opinion than yours, you may grow even stronger biases. Imagine if you I don't know in Italy so is very important. In Milan, we have Inter and Milan that are two the two main football teams here. And practically if you take a people that a person that roots for Milan and put him together with people from the other football team like from with the 10 people from Inter, he will not become more pro in. He will not like Inter more. It probably will fight more and will root even more from its previous team. So the way algorithms choose to put people together, make information flow affect the opinion dynamics. So we don't know anything about how algorithms do that. And often even if you have a very simple let's say reinforcement learning algorithm that is just uh improving engagement over multiple time steps we don't know what is going to happen because the dynamics becomes very hard to understand but we don't know what is the reward. So at least we need to have that to understand what's going on and to help people to have a better life online and have less less negative effects. Sorry this looks I think very general but we have I don't know I think we can talk of this for hours. >> Sure. Yeah. At this point it's almost a product question. How do you deliver this insight and you know show some transparency to the user? And I guess maybe the platforms themselves will have to come up with a good way. But I'm curious from your research, do you have any sense, maybe this is an impossible question, I don't know, but do you have any sense of how users respond when offered the opportunity to see some transparency? Uh like the recommendation of we didn't show you this and here's why. uh how do users generally respond to some feedback like that? No, this is a very good question and I I was thinking just now that you were talking if you remember at the point Twitter was losing a lot of users and they were moving to blue sky and the idea was that blue sky was more transparent in the way it managed its uh recommendations and the same thing is happening to things like masttodon no so some people maybe but they are usually quite aware of recommended dynamics in social media for them important to have a more transparent and safe recommener. It will be nice to see about this. Uh I think with Gregor we did some other studies where we provided for example information on how reliable was the content and we were trying to see if this would help people uh recognize and remember if some information was fair or not. And I don't remember now totally the result. I think uh we found some kind of nudging that made users more aware of how much the information they were seeing was uh reliable. Gateor, you may remember more than me, I think. So I think what we did was we actually had like photos of people and it simulated like the social media environment and the task was to like assign a label to this picture and the participants were provided a AI generated label and another label and they had like to to choose I think if they trust the AI generated label but in this exper experiment I think it turned out that they didn't trust the I generated label too much >> in my I I think uh your work is groundbreaking in certain senses that there are new ideas presented here that maybe aren't yet popular in the recommener system community. Do you have a vision for what the field will look like as they better understand and absorb your insights? >> Well, I think uh on the algorithm side there is there has been a lot of change uh recently. So we were proposing to use more fair reward seniors for reinforcement learning based recommenders already three four years ago. Now this is happening. Uh again my feeling is that a lot of money come from having the algorithms faster and more effective and more safe and aware of the user well-being. So there is not a big change there yet. A lot depends on how much data will be available to the community and also fundings to do research. Nowadays this is a bit changing one way or the other. So I I think people are getting more aware of these issues and the attempt to get more data will surely help other people to play with it because the fact that for example audit you could not use much many data to do audit. In these cases, we can do audit with much less data than what was available before. In most cases, they add to have simulated users interacting with recommenders for a long time. And then we say something about the recommener. In this case, we can just test the recommener without telling anything, without creating fake users and so on. So this approach is a bit more brute force but a bit more machine learning purely machine learning approach and can lead people to do more challenges they find better algorithms and uh you know how machine learning works. If you have a new data set, you have a new interesting problem, people play with this new game, no new toy. >> So I think this may push a bit the community in the direction. I hope so at least. >> Well, I agree with your earlier statement very strongly that probably to date most investments have been engineering investments. Can we make it faster? Maybe can we optimize a very simple metric like click-through rate? Uh but of course that ignores things like fairness. What are the consequences of staying on this path and not taking into consideration some of the concerns you're raising? >> Well, I'm not a social scientist. I'm still an engineer even if I do many many things in every direction. And let's say that at the moment the world is quite crazy know I don't know the news are scary and if you remember there are data showing that polarization increased a lot. There were I think some data from World Bank that were suggesting that the I I did a talk some last year about this with other professors around the world and uh what they are suggesting the biggest fear at the moment are escalation of wars and escalation of polarization that will bring not wars but big societal issues like riots and these kind of things inside each nation if they don't allow you to see other people point of view but keep showing you and pushing you making your opinion more and more focused on just one position and not understanding the other. This contrast with other nations with other people will get stronger. So yeah, probably this makes people make money. People fight online a lot, but uh uh I I don't know on the long run how much this will pay. And so if we want to have a more peaceful and more rational social interaction, we need to be informed about other people opinion. We need to be informed in a way that we can understand it not only as an aggression from the other perspective. No, we don't want to see the other people perspective because in many cases present as something that attacks your state, your opinions and so on. And this is because in many cases the content that gets more engagement is the one that is more annoying is not the right word but you understand what he mean is things that get people more upset and this gets more reactions and is shared more. So this doesn't help interaction and improving social well-being. I think if we don't try to avoid this uh yeah things already went quite bad and I think they can get worse. >> Yeah, you raised an excellent point. I mean what on the surface seems like just a little simple algorithm to surface content can potentially have a profound impact on polarization. Do you have any sense of how we measure that? How can we know to what degree polarization was algorithmically caused? >> The problem is there is very little data on this. Uh so some two years ago I think meta alloed some groups to study on data and even change recommenders during elections but was on a small set of participants and in many cases was really close to the election end. So when you do that in that condition it is too late. It is because most of the people you are contacting already have their minds set and so you cannot see effect of the algorithm. This is something we were trying to do some time ago but the problem of collecting data actively is uh it is hard because there are so many bots online for example. know and there is a limit on how much data you can collect just through API or through scrapping and these kind of things. So the problem is most of the data we got at the time I think were there were a lot of boats votes from advertising things on Twitter and so on. But what we were trying to see what who we were trying to follow were people that were uncertain that we were thinking could follow both directions. For example, in Italy, there were two main party that were apparently fighting and at the end ended up in the government together. This was actually very funny, but uh we were trying to follow these two separate part. What we are saying is that we were trying to collect data on on people that didn't already a fixed mindset, were not already white or black with their opinions. And uh this was hard but I think that's what you can look at when you are trying to understand what is the algorithmic impact because I don't think that algorithms can effectively release which a person that is rooting for one football team to move to the opposite football team. It takes a lot of time. It is very hard. You know there are some works about this with a game called the diplomacy where for example meta was playing with a large language model uh and studying their capability to actually be uh to influence user opinions. So in that case is uh there are controlled studies on this on social media. In many cases, the data we have are not completely clean and they are too close to the end of the decision making process that was lasting for a lot of time. And in many cases what they did was selecting people from that already were for a team and try to see if they will switch and they will show that recommenders will not affect this. But my impression is that you don't have to focus on people that already have an opinion. You have to focus on people that are uncertain to see the impact of the recommener. So, and that is harder to do especially if you are not the owner of the platform. Okay. >> Yep. >> Um I don't know if this enough. >> No, it's good. You rais a great point. A lot of challenges for researchers in looking into these sorts of things. All right. Well, let's see then to wrap up. I'd like to ask each of you, maybe we can start with you, Sabrina. What's next for you? >> First step is get actually get my master degree and then hopefully enter in a PhD program. >> Well, very cool. I think you're off to a good start for that. Gregor, how about you? What's next for you? >> Well, I'm still trying to finish my PhD. I hope that this will maybe work out by the end of next year. I hope so. And then, well, let's see. Maybe I will work for end up working for one of the big social media companies. I don't know. >> I think that would be to their benefit. So, we'll see. And last but not least, Dimmitri, uh what's next for you? >> You know, once you are in academia, many cases are to get out. Well, there are many ideas we are trying to follow about social media also about large language models. So, there are a lot of toys we want to play with and we will do that quite uh quite soon. And yeah for us it's important to collect data about uh recommener system in social media. So we are a call for people to share data and we will share a data website soon where people can upload their data have a tools to anonymize it. >> I think that's the crowdsourced audit we had mentioned earlier. Yes. >> Yes. Yes. >> It's an exciting project. We'll put a link in the show notes for listeners who want to follow up. Thank you. seems like it's a straightforward way to contribute their data and with the absence of data that'd be great for the community to get together and do. >> Yes, thanks. It will be very useful. >> Well, then lastly, a question for all of you. Is there anywhere listeners can follow you and your work online? >> Okay, so actually I have a web page and obviously we have a scholar page. So that's a easy way to follow what we are doing. At the moment my lab is kind of new and we are creating a web page with all the resources and a website with all the resources and the material and this at the moment I cannot give you the art because we are not we didn't publish it yet but I can give you later and if you can add to the description of the show will be great. Thank you. >> Sounds good. We'll do. And not everyone has something, but uh Sabrina or Gregor, any uh like social media or a web page or anything you'd like to include? >> Well, you can um find me on LinkedIn. Happy to connect. And that's basically the only social media platform I'm using. >> Yeah, also LinkedIn for me. It's fine. >> Sounds good. Well, I'll definitely connect with both of you guys there or all three if you're on as well, Dimmitri. >> Thank you. Thank you all so much for taking the time to come on and talk about the project. >> Thanks. It was great and actually quite we got a couple of ideas that we may try. Thank you. Thank you. Yeah.

Original Description

In this episode of Data Skeptic, we explore the challenges of studying social media recommender systems when exposure data isn't accessible. Our guests Sabrina Guidotti, Gregor Donabauer, and Dimitri Ognibene introduce their innovative "recommender neutral user model" for inferring the influence of opaque algorithms.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Data Skeptic · Data Skeptic · 0 of 60

← Previous Next →
1 Data Skeptic book giveaway contest winner selection
Data Skeptic book giveaway contest winner selection
Data Skeptic
2 OpenHouse - Front end and API overview
OpenHouse - Front end and API overview
Data Skeptic
3 OpenHouse Crawling with AWS Lambda
OpenHouse Crawling with AWS Lambda
Data Skeptic
4 [MINI] Logistic Regression on Audio Data
[MINI] Logistic Regression on Audio Data
Data Skeptic
5 Data Provenance and Reproducibility with Pachyderm
Data Provenance and Reproducibility with Pachyderm
Data Skeptic
6 [MINI] Primer on Deep Learning
[MINI] Primer on Deep Learning
Data Skeptic
7 Big Data Tools and Trends
Big Data Tools and Trends
Data Skeptic
8 [MINI] Automated Feature Engineering
[MINI] Automated Feature Engineering
Data Skeptic
9 The Data Refuge Project
The Data Refuge Project
Data Skeptic
10 [MINI] The Perceptron
[MINI] The Perceptron
Data Skeptic
11 [MINI] Feed Forward Neural Networks
[MINI] Feed Forward Neural Networks
Data Skeptic
12 Data Science at Patreon
Data Science at Patreon
Data Skeptic
13 [MINI] Backpropagation
[MINI] Backpropagation
Data Skeptic
14 [MINI] GPU CPU
[MINI] GPU CPU
Data Skeptic
15 OpenHouse
OpenHouse
Data Skeptic
16 [MINI] Generative Adversarial Networks
[MINI] Generative Adversarial Networks
Data Skeptic
17 [MINI] AdaBoost
[MINI] AdaBoost
Data Skeptic
18 [MINI] The Bootstrap
[MINI] The Bootstrap
Data Skeptic
19 [MINI] Dropout
[MINI] Dropout
Data Skeptic
20 [MINI] Gini Coefficients
[MINI] Gini Coefficients
Data Skeptic
21 [MINI] Random Forest
[MINI] Random Forest
Data Skeptic
22 [MINI] Heteroskedasticity
[MINI] Heteroskedasticity
Data Skeptic
23 [MINI] ANOVA
[MINI] ANOVA
Data Skeptic
24 Urban Congestion
Urban Congestion
Data Skeptic
25 [MINI] The CAP Theorem
[MINI] The CAP Theorem
Data Skeptic
26 Unstructured Data for Finance
Unstructured Data for Finance
Data Skeptic
27 Detecting Terrorists with Facial Recognition?
Detecting Terrorists with Facial Recognition?
Data Skeptic
28 Predictive Models on Random Data
Predictive Models on Random Data
Data Skeptic
29 [MINI] Entropy
[MINI] Entropy
Data Skeptic
30 [MINI] F1 Score
[MINI] F1 Score
Data Skeptic
31 Causal Impact
Causal Impact
Data Skeptic
32 Machine Learning on Images with Noisy Human-centric Labels
Machine Learning on Images with Noisy Human-centric Labels
Data Skeptic
33 The Library Problem
The Library Problem
Data Skeptic
34 Stealing Models from the Cloud
Stealing Models from the Cloud
Data Skeptic
35 Data Science at eHarmony
Data Science at eHarmony
Data Skeptic
36 Multiple Comparisons and Conversion Optimization
Multiple Comparisons and Conversion Optimization
Data Skeptic
37 Election Predictions
Election Predictions
Data Skeptic
38 [MINI] Calculating Feature Importance
[MINI] Calculating Feature Importance
Data Skeptic
39 MS Connect Conference
MS Connect Conference
Data Skeptic
40 Music21
Music21
Data Skeptic
41 The Police Data and the Data Driven Justice Initiatives
The Police Data and the Data Driven Justice Initiatives
Data Skeptic
42 Studying Competition and Gender Through Chess
Studying Competition and Gender Through Chess
Data Skeptic
43 [MINI] Goodhart's Law
[MINI] Goodhart's Law
Data Skeptic
44 Trusting Machine Learning Models with LIME
Trusting Machine Learning Models with LIME
Data Skeptic
45 [MINI] Leakage
[MINI] Leakage
Data Skeptic
46 Predictive Policing
Predictive Policing
Data Skeptic
47 Mutli-Agent Diverse Generative Adversarial Networks
Mutli-Agent Diverse Generative Adversarial Networks
Data Skeptic
48 [MINI] Convolutional Neural Networks
[MINI] Convolutional Neural Networks
Data Skeptic
49 Unsupervised Depth Perception
Unsupervised Depth Perception
Data Skeptic
50 [MINI] Max-pooling
[MINI] Max-pooling
Data Skeptic
51 MS Build 2017
MS Build 2017
Data Skeptic
52 Activation Functions
Activation Functions
Data Skeptic
53 Doctor AI
Doctor AI
Data Skeptic
54 [MINI] The Vanishing Gradient
[MINI] The Vanishing Gradient
Data Skeptic
55 CosmosDB
CosmosDB
Data Skeptic
56 Estimating Sheep Pain with Facial Recognition
Estimating Sheep Pain with Facial Recognition
Data Skeptic
57 [MINI] Conditional Independence
[MINI] Conditional Independence
Data Skeptic
58 MINI: Bayesian Belief Networks
MINI: Bayesian Belief Networks
Data Skeptic
59 Project Common Voice
Project Common Voice
Data Skeptic
60 [MINI] Recurrent Neural Networks
[MINI] Recurrent Neural Networks
Data Skeptic

The video explores the challenges of studying social media recommender systems and introduces a new model for inferring the influence of recommenders on user interactions. It discusses the importance of fairness, transparency, and accountability in recommender systems.

Key Takeaways
  1. Collect data from social media platforms
  2. Apply computational models to analyze user behavior
  3. Use graph neural networks to predict user preferences
  4. Evaluate the fairness of recommender systems
  5. Conduct research on the impact of recommenders on user opinions
  6. Apply reinforcement learning techniques to optimize recommender systems
  7. Use crowdsourced audits to collect data on recommender systems
  8. Create a website for people to upload and anonymize their data
💡 The 'recommender neutral user model' can be used to infer the influence of recommenders on user interactions, and fairness, transparency, and accountability are crucial in recommender systems.

Related Reads

📰
Day 29/100 Koko Eating Bananas (Binary Search)
Learn to implement binary search to solve the Koko Eating Bananas problem and improve your algorithmic skills
Medium · Programming
📰
O(N) Manacher's Algorithm with Mirror Boundary Optimization
Learn to optimize palindrome detection using Manacher's Algorithm with mirror boundary optimization, reducing time complexity to O(N)
Dev.to · Dipaditya Das
📰
Building a Power Grid Inside Minecraft with BFS Algorithms
Learn to build a power grid inside Minecraft using BFS algorithms and understand its relevance to cloud services
Dev.to · Carlos Cortez 🇵🇪 [AWS Hero]
📰
The Run-Length Encoding Trick: How Simple Strings Get Compressed
Learn how Run-Length Encoding (RLE) compresses simple strings, a fundamental technique in programming and data compression, and apply it to optimize storage and transmission of data
Medium · Programming
Up next
Stump Grinder Carbide Wheel Grinds Hardwood To Chips
Innoforge Studio
Watch →