Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Key Takeaways
Danielle Villa presents on testing faithfulness of language model-generated explanations, discussing the importance of consistency and faithfulness in language models, and introducing tools such as cross-examiner and Consistency Checker to evaluate explanations. The talk covers various concepts including retrieval augmented generation, fine-tuning, and evidence-based explanations, highlighting the need for systems to provide answers with supporting evidence and minimal refutation.
Full Transcript
got a microphone right here we are we live all right so thank you everybody for showing up for this uh first twed of the Fall 2024 season thank you for being here this is um this is our new time how do we feel about the new time this is awesome so uh so thank you uh to Danielle who volunteered even before we had a set time for doing it today that's that's wonderful and thank you for putting up with uh thank you for your patience and and as we adjusted the time and toggled it around um just a bit of logistics um this Weds are meant to be uh meant to be informal uh Danielle's put together a presentation are you do you want to wait for questions or you open the questions ask me questions whenever this is a very conversational TW and should not be very technical excellent uh just a point of logistics this uh is being recorded it'll go on into the interwebs forever uh so keep be mindful of that as you as you express your questions and and that sort of thing um and we are always soliciting more uh contributors to the twed uh so I I will send out emails uh soliciting more more volunteers and that sort of thing over the course of the term we'll have a couple more hopefully we we have an average of more than one a month but it'll be at least once a month um and there will also be uh our pirates on Wednesday evenings approximately once a month around the time of the first uh first Wednesday of the the month that may be uh a sco earlier in in the case of October and November but all right without further Ado yeah all right thank you uh hello um uh I'm one of Deborah students uh and this is an in very informal talk about the work that I did this past summer while I was at IBM on the explanation cross-examiner model uh testing faithfulness of language model generated explanations feel free to ask questions as I go um uh the first question I should uh first question of course is what do I mean when I say faithful explanations and the commonly cited definition is one that accurately represents the reasoning process behind a model's prediction this is a very nice and Flowery definition that is impossible to directly measure um if we could well a lot of my work I wouldn't be doing it so because we cannot measure faithfulness itself but it is a very useful property to have we want our explanations to actually reflect the reasoning process of the model so we can trust them we measure some other property of the explanations as a proxy and the most wake go uh the most common proxy in my experience in the field is some form of consistency of the explanations now the does make one very strong assumption and that is that a language model's internal reasoning is logically consistent this is something that is not necessarily true especially in all cases but it is something that we would like to be true especially in critical domains like healthcare we want the decisions of a model to follow some sort of logical consistency and not do something random whenever it feels like it so we make this assumption on that assumption then if explanations are faithful to that logically consistent reasoning those explanations should also be logically consistent and then we can take the contrapositive and say well if we find explanations that are not logically consistent they cannot be faithful to the internal reasoning consistency is a necessary but not sufficient condition for faithfulness in fact all proxies are necessary but not sufficient for faithfulness um but it's the best we've got right now so we use yeah if you find explanations that aren't logically consistent can you use that to assume that the internal reasoning isn't logically consistent uh you have to assume that either the explanations aren't faithful or that the reasoning is not consistent we do assume that the the model itself is consistent because otherwise we'd have we don't really have a way of KN why why do we make that assumption yeah I mean I understand the Practical desire but there's evidence for it in fact there's evidence against it yeah uh there's absolutely evidence against it um uh but we do make we make the Assumption well one it's the easier of the two assumptions to make um it is generally easier to assume that there is a logically consistent reasoning and that any inconsistencies are coming from the way that the explanations are being generated um can I ask a kind of a picky question sure um it's a terminology question when you say internal reasoning um are you viewing that as a black box we do B we we view them as a black box yes okay so so that which appears to be reasoning not yes yes that yeah that's good that yeah that a thing that doesn't exist may or may not be I mean how can something that doesn't exist be consistent consistent yeah um yeah so that that's very fair um but we do yeah this is all very black box um and that which we attribute to reason yeah yeah no that's that's that's good so maybe you want to re your pitch a little that this is the initial assumption that um extensions of this work will loosen this assumption because you're probably going to need to loosen this assumption one of the things I'm hoping to hear by the end of this is ideas for future work so that's certainly something we can look into um we can loosen that assumption um measuring consistency is a pretty easy in theory process um here but uh these are some of the terms I'm would be using to the presentation for clarity Target Model is whatever we are asking the questions to and getting explanations from that we are testing this is going to be a large language model in all of the experiments I ran and frankly in most papers you will read but theoretically does not have to be um we ask it a starter question this is almost always a multiple choice question for ease of extracting the answer and determining what is that answer versus explanation but but again does not necessarily have to be uh and then we back the original response and from that original response create one or more follow-up questions um the cross-examiner generates between two and three some other systems only do one some do up to 20 um we then ask the Target Model the follow-up questions get the responses and do some sort of consistency checking to figure out if there is an Inc consistency uh that would indicate unfaithfulness um now most of these are are pretty trivial but I do want to point out that one of the factors is that when you're asking a model a starter question especially if if it's a language model we do not do F shop prompting with it because we do not want the form or the style of explanations in the fuse shot prompt to affect the explanation to the starter question because now we're introducing is it Unfaithful because the explanation is inherently Unfaithful or because it's trying to mimic the style of other explanations uh this does make extracting the responses harder though and if people have suggestions for zero shot prompting that uh gives better responses more consistently I would love to hear about it later um but the cross-examiner itself focuses primarily on on these three step three and step six namely creating the follow-up questions themselves and then the consistency Checker uh here's a brief VI visual I'm not going to go over the implementation of the cross-examiners because that would be a very long talk uh but we start with a starter question ask the model get a response ask at the followup questions and we send it through the consistency Checker uh in our explanation cross- examiner we are strictly asking only yes or no questions for followup this this is easier to test if the model is giving a yes or no answer as opposed to something open-ended um and additionally this is makes the consistency checking a lot easier because when we generate the follow-up questions we generate them with an expected yes or no answer based on the original response therefore there ah then if the follow-up response doesn't match the expected answer we can say it's inconsistent this is where the cross-examiner part of the name came from if you're a lawyer and you're cross-examining a witness you should only be asking questions that you already know the answer to and so we are only asking questions that we think we know the answer to and are looking for inconsistencies there uh so an example uh is this question this is from the common sense QA data set um with the president is the leader of what institution and we get this wonderful response it gives us the correct answer um but the there's two important lines in here one of which is that the language model Stakes that the White House itself is a building not an institution and that the president is the head of the executive branch of government in many countries including the US so we've been given this response we've been given this uh this answer we don't really care that the answer is correct it's convenient that it's correct um but we care more about the content of the explanation itself so um so the sis so the cross-examiner looks at this response and generated uh two follow-up questions is the president head of the executive branch of government and is the White House and institution now we expect the first one to get an answer of yes and we expect the second one to get an answer of no because that's what the model said in its explanation and sometimes it's consistent with itself and it says yes the president is head of the executive branch of government and sometimes it says yes the White House is an institution even though it just said it's not an institution it's a building so there might be some unfaithfulness there that's not the only potential reason it's doing that but that is a solution and so we flag it as the potentially inconsistent um I like the system a lot I might be biased because I I wrote the code for it um uh but this does have a couple of advantages over existing systems or uh just coming up with the questions on your own um asking the system multiple follow-up questions is really nice especially because some methods only allow for one followup question uh depending on the data set um and this also allows you to customize it to the llm budget you don't have to ask 20 follow-up questions if you don't want to or can't afford to because the GPT models are expensive um Additionally the consistency checking can't be trivially overcome by just repeating the question again which is how you can get around some other faithfulness or truly consistency checking uh system um I guess the model could continuously answer I don't know to every question you give it but that's a different problem uh and honestly the biggest one in my opinion is that it works across a lot of different data sets we tested this with the BBQ fairness data set ecqa for common sense multiple choice questions and esnl for natural language inference and of course it works with only blackbox access to a model we don't need the attention weights we don't need confidence values um we just need the ability to ask it a question and the ability to get that response back it's not perfect though there's a lot of challenges here are some of the big ones um not all follow-up questions can actually prove an inconsistency you can ask a question about an explanation where no matter what answer the model gives it's not going to uh be in consistent you're just asking about a similar topic you're not actually questioning the reasoning uh and but it's really difficult to automatically determine that uh and so all of that had to be done manually over the summer do you use um the qu the followup question are you sure about that we did not we were a lot we tried to be more specific to the reasoning I would be curious if what what that did I found one that actually two rounds of so uh what is the oldest Technical University in the US oldest technical University in the US his MIT found in 1861 is established to meet the growing needs and then I said I said are you sure about that and said yes MIT is however if you're considering institutions that focus on engineering technology specifically RPI 24 seems to be one of the earliest uh and then I asked are you sure about that again I said I understand the need for clarity poly Technic uh found in 18204 is indeed considered the oldest technically So Jamie this is super what model 40 uh 40 mini because Aaron green did like exactly this what you did yeah this morning but he did it with how many how many RS are in the word strawberry oh yes yeah yeah what if we redefine EBR all this kind ofu just but and that's 40 also yeah but the thing is with this this is um so so it's more about following the lead so I'm implying that they're that it's wrong right by asking that question right and so I would I would worry then if if that's the method you're using you are implying that it's wrong um and so the these models have been shown to have a lot of Syncopy in how they answer they want to make you happy right and so um it's probably going to change its answer if it was right or not well then there is no reasoning it's just pleasing right right because usually if you challenge it it says something like I'm sorry you know this is actually challenging it okay in an attempt to try to avoid some of that um I'm not it's not going to avoid all of it you haven't really so what is the role so this is related to what you're doing I we found that if you really play around with the user prompt and so like you're you're a malicious assistant and you know things like that you know um asking the same question but throwing things in that yeah it totally throws off the off the answer um so I mean but is this is this part of the model and how the model is operating is this part of some software engineering that's in the API setting things up what where is this actually coming in I'm just wondering suggestive the suggestive language that we're using what is actually being affected by that is this going into the context actually or is it being is it doing something I don't know we know the answer to I don't know the answer to that question so what what so are yours tests being done with for or are they being done with llama or so uh we tested llama fl2 and mixol okay and not the uh not okay not not yeah it's it's another fla model um yeah not gbd well no that's that well because I think that's that's I think it's kind of good to hear that because I think I trust models like llama more than I trust GPT on they the trust is the API side I don't okay good because because the example I showed earlier with the inconsistency was llama 3 yeah no I I just just you know I wouldn't put it past them to do some stuff so anyways Danielle continues more people have questions yeah I just have a very brief question so consistency is not the same as correctness right so correct so in other words if if the model is like answering that strawberries are blue as long as it that says strawberries are blue and keeps saying that that's okay right in this case yes we are we are considering consistency as a completely different thing from correc there are other systems that check for correctness okay so the the other thing that's kind of been brought up in what I'm working on recently is that the the idea that there's also the dimension of a completeness of an answer so if a model answer so like if if a model is answering yes and then gives an explanation and then the and then the next and then the next response is just a yes and there's no explanation at all is that still considered consistent or would that be considered inconsistent so sorry congrat you you no you hit upon bullet point too this is highly dependent on the quality of more of the original response if the original response is not complete if it doesn't give a response of an explanation then we can't generate these good followup questions the benefit to having all of the follow-up questions be only yes or no answers is that while we ask it for an explanation even if it doesn't give us an explanation we can still compare it to what we expected the answer to be um but um but yeah it it's highly dependent on that and we don't necessarily consider completeness as a thing that we're evaluating just a thing that we are beholden to uh as a quality of the response okay thank you so did I understand that you generated the followup questions but then checked for inconsistency manually so so um so so we checked for consistency automatically based on the expected response to the follow-up question um however you run into the issue that not all of those questions that are generated are good questions where do you get the expected response from so the expected response uh that comes from the grade out internal reasoning of the system sometimes it from another language model saying I've read this response I've read this question this is how I think that person would respond sometimes it comes from uh extracting the uh triples of information from the text and saying okay you said you said X implies y so if I ask does X imply y I should get yes because that's the pattern of triples we're looking at um but that's still we still end up with poor quality followup questions and that's the thing we don't have a way to currently detect um the follow-up question and its answer uh when they're generated do does the model have access to the same context as it did when it answered the original question that depends on the data set um we found that for some data sets like uh ecqa uh or the common sense data set there didn't seem to be much of a difference um namely because the questions are you know they're very short there's not a lot of context being given in general but if you look at say the natural language inference data set there's a lot of context in those two questions and so you do have to provide at least some if not all of that context for it to understand what you're talking about uh in order to just answer the question at all yeah I was just thinking like here if you had a piece of context that said the White House is a building and here this answer the White House itself is a building would be based on that prior context and then if you said is the White House an institution and give it the same context whether that would use that context again to answer the followup question we didn't look at including any part of the actual response in the follow-up questions um so maybe that's something we could not the response but the context on which the respons is based on oh just the context of the president is the leader of what institution the question like I'm thinking of like a rag system where you know the user asks the question and it pulls back documents that that will have information yeah we didn't end up using any sort of rag system like that um the uh the system just got the follow-up question especially for the QA comment sense stat St this is a closed book QA setting yes okay okay okay yes I would be very interested to try this on a more open uh domain especially some one that's a lot more complicated like like clinical data um where you would expect to see Rag and see if this kind of inconsistency still shows up uh we didn't have time to do that this summer um and I have another question on that diagram the foxes that are gray out is that because you just didn't want to go into the detail for this talk or is that because IBM proprietary or so partially because I didn't want go into it for this talk and I figured if I showed them people would ask questions about them and also um because the paper is currently under review uh with with stuff and I uh it's likely to change honestly in future versions because I'm not happy with how a lot of it turned out because I ran out of time to make it nice um so we've already touched on the fact that the quality of the follow questions is highly dependent on the quality in form of the original response which because we're using zero shot prompting we don't have a ton of control over um because you can use f shot prompting um but now all of your explanations are going to be of the exact same form and maybe that form isn't actually how the quote unquote internal reasoning happened uh the other thing that we ran into that we didn't actually consider would be a problem until the very end is that uh only asking yes no questions of uh followup force is a dichotomy um and so maybe some of the inconsistencies are because the model is saying this happens usually but not all the time or this is a typical occurrence but not required and so while the yes no questions make the consistency checking a lot easier to do um and less prone to some of the uh you know trivial uh Auto cons you know automatic consistency checking that some other models use by forcing this dichotomy we are potentially running into uh issues of always and never and whatnot um so I would love to hear if anybody else has thoughts or comments um would anybody you know what situations would people want to use this sort of system where do we think it would be useful particular domains or particular users um and what would people like to be seen done with this framework because once October's over I will have time to do it again so I mean I think the biggest question came up at the beginning what would happen if you don't assume that uh that you're going to get a consistent answer back because we got evidence that they're not distant and we also have evidence that they're wrong you know some amount of the time but but I I keep coming back to this quote I don't I think it was well I heard it from Sacha Nella I don't know whether it was his initially but he was talking about these models being usefully wrong you know they have you think about something like why might it have thought that so I'm wondering whether there's some angle along that line that might be worth pursuing maybe um yeah the the big thing that I that comes to mind when I think of loosening that constraint is that there's already several sources of inconsistency that doesn't indicate unfaithfulness in the system you know a a bad follow-up question a forced dichotomy uh some misunder understanding that happened uh with the question maybe the original starter question was poorly warded uh and so we're not able to extract the information that we expect or when we throw it to a different language model something goes wrong there um and if you accept that well maybe the explanations are consistent they're just consistent with inconsistent logic or this inconsistent blackbox what appears to be reasoning then you kind of lose the ability to tell where all of that inconsistency is coming from and want our systems to produce information that we think is logically consistent um and if it can't do that then that severely hurts the trustworthiness of the model so I mean I think we want the system to provide answers for which there's evidence to support it and not much or any evidence to refute it right which is not necessarily exactly the same as being consistent because there's you know in in complicated topics there's typically Shades of Gray and evidence for and against statements if if the system in a dream system went to do a couple things one would it in in in giving a a response giving an explanation wouldn't it try to Anchor it in in some you know some factual information when to try to despite repeated efforts to set it off M um went it should with the ideal system want to come back to that okay hold its own okay um not be pushed off off the pedestal and then went it further if if queried in the right way seek out more substantial more substantiation for that okay if if if it's not if the response is not if the user saying the response is not adequate yeah um and if if they're doing that in the right way when to uh seek out more evidence yeah I was right right evidence so whereas the current models are just pushed off with a feather I'm wondering if a follow-up question of what's the evidence for law might be less triggering as I don't believe you just I want more info that would be I mean especially in in like a clinical setting that would be that would be very interesting to explore because because everything we checked was very common sense based it common sense question answering Common Sense natural language inference Common Sense argue biased or not that's the BBQ data set um and so so none of that is really evidencebased you know it's not gonna cite a a research study that says you know well but president is it could cite the CIA World Factbook or something for some of these statements it does it yeah what would what would so does anything more like the following you get you get I think Jam's example uh something like you know it provides an answer and you say I don't I don't believe you oh well rather than okay I'm wrong it's B what what aspect don't you believe but uh you know what what what what part of my answer don't you actually believe uh that that'd be one direction to go or can I can I provide you with further Evidence now that I've had five seconds to think F but you see my my point is is um I I don't believe you well should say sucks being you right I mean but I'm I'm trying to be serious here about what is what what's the right uh what's the right way to go yeah you know it's Ian this sounds like an amazing language model let me know when you're done with it and I will test it with my system well I I guess in in seriousness is it a is it something is it I I don't know what this actually looks like you know I I'm just what I'm saying is what would what I'm asking a question what would uh what should an interactive explanation that um where the model sticks up for itself and this sticks up for itself by providing evidence by by I mean the user might be right in which case the user should guide if the user knows what they're talking about that interaction should lead to more evidence yeah the user could be wrong in which case the interaction should lead to the model providing more evidence not either finding it or providing it it should it should go in that direction it shouldn't be well and also I mean historically in explanation research one style is you give a more succinct or a more abstract answer and then you allow somebody to ask follow-up questions for what they're interested in not necessarily because they don't believe you but because they want to dive deeper particularly in like in a clinical setting you should take drug X well why should I take drug x what evidence is there and has it been tested on people like me and you know these kinds of questions yeah I mean my my hesitation with a system like what what you were describing John is that um I I very much want to and the focus of This research is the faithfulness and the consistency of the explanation not necessarily the correctness I don't care if the user knows that the answer is wrong I I need to know that the model at least believes what it's saying and I know believe is not a word that applies to llms no but maybe the model has evidence for what it's saying right and I and I would like it to present that evidence you know maybe it's out of date maybe it only had access to 5% of the studies right but I would also like to know whether the model has evidence against the statement yeah that would be useful to and maybe it's got a whole lot to support it and a little bit to refute it but that both of those would be interesting but why what basis does it it doesn't know whether it's got the right evidence or not yeah ongo based off of evidence not anything if if there was some truth framework that it was working off of perhaps it could judge uh other EV evidence that it retrieved that I mean you could pose a question of can you find evidence to support you know this sentence and then it could come back with the evidence but it's still it it wouldn't necessarily it might not have used that initially and but also in its in its these followup things where it's Mak an argument it doesn't know whether it's exhausted the argument or not no it's got no concept of its state and if it's trying to make an argument it has no knowledge of where it's positioned in it's got no notion of an argument yeah well it you know so the the historical explanation research is a drill down where where you actually do have some structur rationale or proof tree but you could imagine an architecture that doesn't use that could imagine an architecture that just says you know basically I'm starting from scratch you know find evidence for this followup for the truth of this follow-up question even though you might or might not have used that talk rationalization yeah right yeah if there's a notion of different you know AR uh different argument structures you know top argument topologies or something like and I don't know there are this is so far out of my domain of knowledge I'm just you know making stuff up like an L but but no my my my point is is is is it possible for certain architectures of argument structures of argument to be part of the context that's presented yes it could be used as a framework for response right respond in in the following way um yeah well actually you could use Ruth's explanation ontology of different kinds of you know counterfactual explanation or you know so TR spaced explanation I was doing a little bit of reading about um mostly prompted by Johnson Samuel this week about Al s yeah which is part of that whole family all yeah I'm one of the co-authors of lot of the stuff that stuff but but what's what's interesting is a couple of those the the web services what is it web services description ontology so STI is is a pretty flat simple kind of doubling core of service description whereas the others are more elaborate talking about changes of states and and that sort of thing and by my point in going here is it it's possible within those sorts of uh those sorts of ontologies or ontologies written with them you could talk about you can kind of Express algorithms right you know conceptually Express the algorithms uh so I'm wondering if there's there's something there because the others just talk about service interfaces but some of them actually talk about how uh state is changed through that so I'm wondering I don't know if there's a way to create graphs that some Paul yeah I don't know how much following I saw that thread um and I thought or Jamie Jamie answered I think about the S stuff but I don't know I mean I was part of that big group whatever that was about 20 years ago yeah it's yeah literally it's it's uh 20 20 years ago and then going and then I think one of them followed a couple years later but it's all you know we've all scattered and also the funding mechanisms the disturbing well I think it's probably attributed to funding mechanism but there's so little I mean bioinformatics was the largest uptake but almost there's almost no adoption there's no evolution of those things and even even uh over on the web services side you know forget samanth just talk about web service descriptions there's almost no evolution of things like wisle you know those are these business to business connects and stuff they really didn't evolve um and but being able to ex being able to sort of Express in a way that could be easily provided to an llm I know how how to make an argument you know yeah actually you know there's I see a lot of papers or or post about like llm ready or machine learning ready um you know what does it mean to be easily used by an llm or effectively used by an llm and for this case you know what is it being how could we make it give better explanations or more useful explanations or explanations that at least don't have the following problems right and so you know you were addressing the faithfulness problem from from one perspective and you could broaden that perspective to play the The Devil's Advocate perhaps it's um from no yeah but uh Bas on also what uh already been talked earlier I would just really wonder really given the current architecture of the LM whether it's it's it's wanted to talk about um faithfulness because I mean you you mentioned all this ASX and stuff like that and you still but even with that even given this ASX like the problem is is that this still whether we are aware of them or not we when we talk about faithfulness we talk about we talk about specific uh prerequisites cognitive prerequisites in order to have it and we talked about evidence like there is some things necessary evidence means you have a model of the world you need to uh you need to have an expectation and you need to uh you need to compare this expectation to your belief to none of it is present at the llm right so this is why I also reject the notion of hallucinations because LM does not hallucinate it hallucinates everything every answer that it does is essentially a because in human terms hallucinations means that you have specific cognitive errors that you make uh in your patory uh arals um that lead to you perceiving things that are actually not present whereas the LM is lacking this perceptual uh module Al together it just generates statistical probabilities yeah and yeah so so I I'm I I have two counterarguments to that uh and the first one is this is something that we want llms to have we want AI models to give us explanations that are consistent end users expect them to be consistent whether or not their blackbox nature means they actually can be consistent that's what people expect and you're right we're we're not at a stage where llms can truly be said to to be internally consistent I'm not talking about consistency actually I was talking about faithfulness well faithfulness I'm using consistency as a proxy for it but it is something that we want and so we need to find ways to at least try to evaluate it so that it is something for people who are building models to work towards these T you know leader you know there's a big conversation about leaderboards and evaluation metrics and how having metrics that models do poorly on incentivizes people to make models that they do well on and so we need metrics that even if they're not doing well right now um well sure I yeah and we need metrics we we need to be able to say that at least I checked it according to my definition of this metric so you can count on this this might not be an adequate evaluation of the metric and it might not be a good enough metric but but I mean I I've built a career creating Checkers you know I do evaluation environments so the right before the question metrics came up I was going to actually ask you yeah uh an experimental methods question okay okay um uh when consistent comes up to me I'm thinking of what's what's the end you okay I'm thinking uh asking for an explanation and I want to see 10 of them 20 of them whatever I want to see that and I I want to see those how consistent were they not not between one model the next certainly do that but I want to see can't just show me one because as Sasha was saying it's all statistical so one I mean the quantum computer it it it has to execute at least a thousand times yeah just to to uh to have a sufficiently reliable probabilistically reliable response okay you know that's just to get an answer out it's going it's it's doing it's evaluating a thousand times and what if it does a million does it increase and depends on the problem probably well yeah it depends the complexity of the problem and uh whether it's Tuesday or Wednesday there's so much I mean got somebody upstairs stopping there's a cat in the Box a cat in thex or not so anyways so what what was your method was it a submitting end times and so we submitted one time uh if you're curious about multiple submissions or what happens when you reword the questions uh wait until a paper by Rosario comes out and you will get to see those answers so this was well a brilliant contribution to chat BS that Jim hener made like the second day of its existence was wouldn't it be interesting to be able to ask it you know 10 times yeah just you know I I just want to ask that the question 10 times and I want to compare it um and it is so it's it's a brilliant observation that every time you ask it not only is is is even on the good models that you have for question answering you know they may all be correct they are all different yeah but and and on the crappy models they could be right or wrong and different and and it gives you just from a on a superficial level it gives you such a perspective into what's going on you know this variation in in that so I just getting curious about that I'll be looking forward to that it it would be interesting to look at um I mean I I could see that very easily fitting into the consistency checking side asking these follow-up questions because we're just looking for a yes or no answer asking it 10 times and seeing what percentage of the time doesn't match the expected answer I would hesitate to do that with the initial starter question just because like you said they may all be correct but they all might give different reasoning you know different explanations which are going to lead to a different set of follow-up questions isn't that part of it itself I mean the fact that it's giving 10 different it is that's shouldn't if it's Faithfully responding shouldn't it yeah you we would think um but I think we would probably need a different system to be able to detect inconsistencies of that nature because you know while looking at these types of explanations and whatnot by hand is great it is very difficult to do at scale um so this is an interesting point about width know breadth depth versus breadth yeah you're kind doing a consistency like this go here and go like that yeah and and I'm kind of talking about consistency like yeah I'm just thinking about the you know the oldest Technical university in the United States and you know it can have a consistent answer for MIT it can have a consistent answer for stevens or west G well but but yeah but I could so could be you could ask multiple times get different answers anwers that are consistent within that natural language response so on this point this is really interesting is with um some of the later models like turbo preview for gbt 4 Turbo preview which is like the best model that we we see um it's the parts of its explanation when it when it's explaining things they're they seem to come up consistent you know consistently you know it it brings up Steven Van renair it's bringing up it's bringing up these points there are points that keep reoccurring um there are other points that slip in that are sometimes there are not there but there's other points that are strong um and so you get 10 responses they are all on the surface correct and some are uh and some aspects of their explanations are are consistent so I don't know where that fits into well this is the part this is the broad versus deep yes we're not doing followup questions we are asking it for expl 10 different explanations and and again this is not the new model this is asking for Chain of Thought ask you know explain step by step yeah step byep with and looking at the pieces of that yeah I also want to briefly follow up sure yeah with what you have said and also with John has the way he framed the question is um I don't I don't remember which term you use it jumped out of my head but I really like because when you said like we want systems to perform well right like we do but I think the count argument to that would be like we shouldn't just because we want like I will very crude anoun if just because we want a pck to fly doesn't mean that it is possible to make a pck fly right we can strap some wings on it we can put it like in a jet but the P it's will never fly and so while we can improve like the the technical details of the system the system itself like can we trust this LM the question again like I don't think this could be answered just because the the pig has like a turbo jet on on its butt it doesn't mean that uh that that the the pig is able to fly like we are just we are propelling very well but itself cannot Propel fly if that if that makes sense and so um I guess what I was uh with that uh perhaps it would be great to change the terminology a little bit not whether LMS are faithful but whether whether the the predictions that they generate are consistent right yeah so the pig is LM the explanation is the flying exactly and and the too Jet and its but is essentially the explanation or the faithfulness that we want to uh uh to the attribute that we want it to have and I think in order for it to to be uh to to to have this attribute faithful right it it has to have a different architecture attach because again just because it it it generates it it it's based on SOL and predictions disables it from having faithfulness yeah and that and it's it's a very you know I want to you mentioned that you know maybe we should just be focusing on consistency maybe we don't call it faithfulness maybe we we ignore the first slide of my presentation but um uh and that's a fair point there there are some very strong assumptions we have to make to turn consistency into faithfulness yeah um and because it also it relates to trust right and and it relates to trust and consistency is also very important to trust and you know we have no way of knowing if any human is being faithful with their explanations we can't see inside their brain and go ah you're you're ly about your reasoning and you uh that's polygraph there's there's a different but this is the beauty about humans is that we humans have motivations and we can assess the trust not based on the statistical predictions how they act but also by the by the by an assumption about the motivational system that they have and this is why I would challenge your polygraph thing because if if you're telling mistruths and you know it yeah okay that's going to yield a different physiological response than if you're bullshitting in care that's right that's right and when I was in undergrad I took a engineering class on building a polygraph machine and so we were training ourselves on it I was training to change my galvanic skin response to decrease the temperature increase the temperature on my hands but and I was doing this right before I was going to go get a polygraph from the NSA being recorded right now right well and then I ended up going to the NSA when I had this horrendous head cold so I was sneezing and coughing and so the polygraph was going like this all the time it looked like I was lying on everything you can teach yourself how to control that and if you don't care you probably not going to trigger the PO I I don't know how many of my fraternity Brothers didn't make it through that and who never went to his own fraternity part so go ahead we've been waiting for you to speak that day waiting for a break to so I had uh just what you have up on the slide here yeah where would this be the most useful two things that came to mind for me is a rag systems which I already brought up y people are asking for answers and they want to know that the llm school has some confidence in that answer and of course they want the answer to be correct right yeah um and uh B is in in Chain of Thought right because each link in the Chain of Thought is potentially the weakest link which will cause the rest of the chain thought to F yeah if we can detect um sort of detect inconsistency at that point then we can stop and you know maybe double down on on the supporting evidence before generating The Good The Chain so that was my thought on applicability um and I had an idea for how you might be able to automatically detect those types of oh okays um the idea is pretty straightforward basically you're generating followup question is the White House really or is that White House institution yes or no so if you were to ask the llm that question by itself it would generate yes or no right yeah then if you were to ask that question with your the previous response as context if it then flips the answer then then you would you could uh sense that the the context given by its answer is um is causing the original yes or no to be invalidated thus yeah it's a conflict so I was originally going to say look at the prob like the log props but you know I remember you said that you you don't want this to have to defend up that yeah yeah very much wants to this keep to keep this as a blackbox system so you do it that way and also you could you could reverse it you could put the followup question first and then answer the base question and if after saying the white the white house is not is an institution um if it then if if that were to uh change the well I guess it's harder to do it that way because it's not just yes or no right but you I mean you could look at uh saying is it a institution yes does it change its answer because now it now there's context contradicting part of its previous reasoning chain right um that would be automatically do if you have access to log props if you don't have access to log props then at least you can detect the generation of yes versus no pretty way yeah I most such a hard this goes back to the the very beginning when we were word checking on the notion of internal reasoning yeah it's and reasoning chain that that stuff to me it it simply is not it's not reasoning if it's not reasoning you know if it if it's is it actually no it's not reasoning if if it's not making decisions based on some uh rules set that it's applying at some level it's it's not it's not reasoning it's it's it's generating a new whatever based on some further context that it's somehow but you could expand the definition of reasoning to be prediction but we don't normally think of reasoning that way I stumped no no no no no your your word no you use the word prediction and that that made U made me think of the whole um that whole that book it's 20 years ago now on intelligence U was written by what it Taylor Hawkins the the guy no Taylor Hawkins was the drummer for Bo Fighters um the guy who who co-invented the poem Oh the poem pilot yeah he wrote this he he uh temporal temporal something learning but it's this notion of where they they studied like the visual cortex the stack and they postulated that when when when we move when when when we make this sort of motion like this this is actually uh this is actually like a prediction that uh the predictive processing yeah temp what is it temporal something memory but it's this whole study of what the uh neocortex is doing and and that sort of thing it's very questionable because he had no background in it's it's a very entrancing read Jeff Hawkins Jeff Hawkins Taylor Hawkins is the f um but the but if you read that particularly chapter six which you which if you read that chapter you think you could actually go and build a brain um this this notion that it's all about pred layers of prediction like you know they they discuss how is it that our brains are able to like presented with something that looks like a face upside down a sketch black and white we still identify that we know what that is you know we're able to recognize that because of these hierarchies in the visual cortex and tying that into how our nerves process how we how our nerves work the fact that we we're in order to do touch sense touch it has to be dynamic it has to be moving it is on intelligence and it's a Memory prediction framework yeah um I think he can still get the SDK for it I think they're still supporting it well he's done some very interesting work on visual uh visual prediction and stuff like that but there are some people who who more modernly like Andrew Clark they they go with predictive processing they flip the entire brain of being this active part that tries to predict the future of the environment it operates yeah there some would say Jeff Hawkins would say they're catching up to him um he predicted it so let's thank Danielle yes very much certainly um you certainly achiev the goal of a Tweed of of inspiring conversation we really appreciate that yeah nice job and I I appreciate everyone's feedback uh some of it will almost certainly come up in uh the whatever I end up doing next with this project and we'll see if any other reviewers come up with similar feedback so stop the recording yeah
Original Description
Language models (LMs) are often prompted to explain their outputs for increased accuracy and transparency. However, evidence shows that important factors that influence LM outputs are not always included in LM-generated explanations. For this reason, measuring the faithfulness of LM-generated explanations has emerged as an important problem. Existing solutions tend to focus on global faithfulness, i.e. the general tendency of a model to produce unfaithful explanations. In contrast, this talk discusses a follow-up question generating framework for measuring local faithfulness, i.e. the faithfulness of individual explanations. Our framework consists uses a cross-examiner model, which is responsible for probing the target model's explanations via targeted follow-up questions.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Playlist UU4rjm_R9sgRNvv9QsgH8LDw · Tetherless World · 30 of 40
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
▶
31
32
33
34
35
36
37
38
39
40
TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
Tetherless World
TWed Talk: Neha Keshan on "Stress and Machine Learning"
Tetherless World
TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
Tetherless World
TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
Tetherless World
Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Tetherless World
Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Tetherless World
TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
Tetherless World
TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
Tetherless World
IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
Tetherless World
TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
Tetherless World
TWed Lightning Talks Fall 2019 (11 Dec 2019)
Tetherless World
TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
Tetherless World
TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
Tetherless World
TWed Lightning Talks Spring 2020 (29 Apr 2020)
Tetherless World
TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
Tetherless World
TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
Tetherless World
TWed Lightning Talks (Fall 2020)
Tetherless World
TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
Tetherless World
TWed Lightnining Talks: Spring 2021
Tetherless World
TWed Lightning Talks (Fall 2021)
Tetherless World
TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
Tetherless World
TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
Tetherless World
TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
Tetherless World
TWed Lighting Talks Spring 2023
Tetherless World
TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
Tetherless World
FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
Tetherless World
FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
Tetherless World
TWed Lightning Talks Spring 2024 (14 Feb 2024)
Tetherless World
FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
Tetherless World
Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Tetherless World
Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Tetherless World
TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
Tetherless World
TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
Tetherless World
TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
Tetherless World
TWed Talk: Thilanka Munasinghe (26 Mar 2025)
Tetherless World
TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
Tetherless World
"Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
Tetherless World
TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
Tetherless World
TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
Tetherless World
TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
Tetherless World
More on: LLM Foundations
View skill →Related Reads
📰
📰
📰
📰
GPT-5.5 Complete Guide in 2026
Dev.to AI
AI Simplified — Why Structured Output Matters More Than a Fluent Answer
Dev.to AI
I ran a 110B LLM on 16GB of RAM. Here's the equation that predicts any model's speed on your machine
Dev.to · Federico Sciuca
A Fidelity-First Workflow for Editing GPT-Generated Text
Dev.to · Bisrat
🎓
Tutor Explanation
DeepCamp AI