TWed Lightning Talks (Fall 2021)

Tetherless World · Beginner ·📄 Research Papers Explained ·4y ago

Key Takeaways

The TWed Lightning Talks cover various research topics including table question answering, natural language understanding, fairness metrics, and knowledge graphs, utilizing techniques such as retrieval augmented generation and fine-tuning, and tools like ontology, RDF, and semantic data dictionaries.

Full Transcript

all right i'm going to look at the camera so good evening everybody and welcome to the fall 2021 twin lightning talks this is our uh our regular end of semester um celebration of everybody's work the rules for lightning talks are simple two minutes talking about whatever you want to talk we just ask that you come up here we're going to change the camera angle here in a second to point to the podium you just stand kind of center over there and i guess what we should do is let's uh honor our remote guest first fake favorite would you like to go first and if you could please turn on your video for us that would be wonderful oh can you me are you just here to listen um can you guys see me i'm not really sure we see you now we see you now yes okay that's great uh so yeah um pretty much started um two years um i've been working on the table question answering task um we started with um given the question and the table searching for the um the right cell to answer the question um and then we move forward with another um like sap task that's retrieving the table from a large table corpus for given questions so now basically we're building an end-to-end system that um can retrieve the tables plus answer the question um so um one demo paper has been submitted to um acl and another paper is now under submission similar technique a similar task different techniques um with the um with the um acl short paper other things also we are now moving forward to the uh multi-model table question answering means that the corpus not only has tables but also have passages not long like really long documents but a a trunk of text so um there's two data sets that's designed um specifically for the multi model table question answering one is the hyper qa data set and the other is the um ott qa data sets which are now very popular for this community and if anyone's interested um or um or have any ideas or want to use this type of data sets um in their research um please reach forward that i can um have a discussion with you that's pretty much everything all right thank you thank you i i apologize i didn't say you know when when you begin your your lightning talk you know obviously say here we are but also say who you're working with um and if you want how long you've been around here that sort of thing sasha i won't make you go next because you're new to the group uh and new to lightning talk so i'll let you listen in on a few lightning talks to kind of get the feel unless you you're ready to go but i'll i won't put you on the spot yet um we actually have a few people who actually signed up um and sola is next solar would you like to go so you want me up at the podium yeah i'll make it go to the pre-programmed location um you want me to get rid of jim and bernie are they okay okay so uh hello my name is solo shirai and i've been working in deborah for uh i'm on my third year now so that's where i'm at for working towards my phd and so today what i would like to talk about is uh this idea of context as it relates to semantic web technologies and particularly knowledge graphs the idea of context as we use it colloquially clearly is of some level of importance to help kind of understand either the environment we're in or the information that we're working with but in actual research related to the semantic web technologies it's often used very loosely or on the other hand it's defined but it's defined in a very restricted way that's not really usable by other systems or it's not really easily extended and so what i've been exploring is kind of how might we be able to define this idea of context and represent it in some sort of way that works nicely together with this magic web and uh related to this as i'd also like to share two ideas that i've been thinking of related to this idea of context uh the first is the difference between static and dynamic context so static context is what i think of for information that you can kind of define beforehand that perhaps defines the environment or a particular system that you're working in that stays unchanged over time on the other hand the idea of dynamic context perhaps could be useful to consider information that is contextually relevant that changes over say the same run of a particular system or as information that you have available and the second idea about context that i'd like to think about is related to useful and interesting information from the context instead of just grabbing everything that's connected throughout a knowledge graph and saying oh this is context here you go do whatever you want with it maybe we also should be trying to actually identify which pieces of this connected information through the knowledge graph is actually useful for performing whatever task you're doing and what is actually interesting what is maybe surprising to people what is interesting in terms of actually helping to understand the information we have available a little bit so uh that's all i have to share for now and it's very much a work in progress in terms of how to actually do those things like defining and representing context but uh hopefully i'll have some interesting updates about this line of work down the line thank you all right thank you so much jim i hope you don't mind the yeah yeah i noticed it [Music] but that's okay you might still have a job tomorrow if it's not there by the end of the meeting so uh good evening everyone so uh i'm jason i'm a phd student with working with professor deborah mcguinness so i'm going to do my candidacy talk next friday i will send out a formal announcement so uh so today i'm going to talk a little bit about my thesis work uh which the title of it is uh natural language understanding with semantic parsing so so natural language understanding is a sub field of natural language processing so nlu for shots so the main task of this topic is to transform a natural language test into a machine understanding a machine understandable format so so the so the core task of nlu is is semantic parsing which is also uh trying to convert the natural language test into a uh logical forms uh specifically uh i mean semantic representations so those representations can include uh logic logical forms or meaning representations or or some actual executable programs so so we are trying to uh we evaluate uh our approach in uh three sub areas so the first sub area is is a large graph question answering specifically we are trying to answer uh natural language questions using the knowledge graphs so we are trying to apply the semantic parsing techniques to generate some intermediate representations like semantic choreograph to help the generation of the final executable sparkle queries so the second sub areas we are going to tackle is the math word uh problem solving what it means is that given a description of a method math word problem we are trying to use one of the semantic parsing the parses called amr abstract meaning representation to help the generation of a mathematical expression that solves the mathematical problem so the third sub-area we are trying to evaluate on is a large prediction knowledge triple prediction task for a test based games called jericho work so all these subfields are very important um natural language understanding uh some areas in the nlp field so the the connection between these two these three some areas is that we're trying to either uh using some intermediate mini representations or some existing meaning representations like amr to help uh solving the task so um yeah so if you are interested in the my face is work but please come to my candidacy talk next friday so see you good luck with that thank you [Applause] all right so next we have jade oh don't worry it's a huge number we have to cram don't worry about it hi so my name is jade i'm working with uh deborah mcginnis um so i'm a second year phd student uh and i'll be talking about the recent work i've been doing with uh the fairness metrics ontology so this was a project which we originally started it uh over the summer as part of an ibm externship we continued working on it uh over the fall semester and the basic idea is we want to be able to capture all of the uh fairness metrics and techniques that go into assessing the fairness of a machine learning model so a lot of these times someone is like okay machine learning models are unfair we know that's a huge deal but how do you actually analyze the model how do you actually tell when it's unfair and there's not one definition of fairness there's many and in a lot of cases they're mutually exclusive so you can't have more than one definition of fairness be satisfied you have to a lot of work goes into picking what the right fairness definition you want to use is or figuring out what specific kinds of bias you want to address and we figured that all this information can be condensed into a nicely packaged uh in the form of an ontology so that was what we did over the summer and then more recently we've started working on actually using the scientology uh together with wyatt so i've been instantiating some wise representation in order to actually represent um machine learning model evaluation information and say okay this model uploaded at this time trained on this data set has these fairness metrics and then you can use y is of course very providence id you can track over time like okay these additional versions have these changes in fairness and okay these you know other data sets have these different results when you look at their at the kinds of changes that they make to fairness and so what we are currently building towards is ideally a way to just analyze the fairness of some machine learning model so i've been working with some uh so i've been working with heels so i've been working with um karen is another student i've been working with and he has a synthetic data set model which generates synthetic health data sets and so the current plan is to use this as an example um to see how well does this wise based system work so that's all been going on for past couple of months and hopefully we'll be pretty close to producing something interesting okay all right thank you so good hi everyone i'm severe rasheed i'm technically on my seventh year at rpi but this is my sixth year at the lab i i work with deborah mcginnis i'm a phd candidate here so um my current research relates to different forms of reasoning um uh well i'll talk to you guys uh about today is some earlier work um on semantic data dictionaries but i'll um kind of conclude with some future directions that we're focusing on so what is a semantic data dictionary so this is a an approach developed by our lab for uh annotating and transforming tabular data into a rdf format so the semantic data dictionary is made up of several tables so you have the info sheet which links all the tables together and allows you to specify metadata about the data set you have the code book which is used to categorize categorical data and numerical values map them to classes or resources you have the timeline which allows you to kind of more explicitly mapped time series data um and you have code mappings which allows you to create shortcuts within your annotations so if you have like units instead of looking up for the uri for meters you can just go m and then have all the units mapped already and then finally you have the dictionary mapping uh table which is kind of like the brennan butter of the semantic data dictionary which allows you to actually create mappings for both the columns within the data site as well as um implicit objects elicited by the data set or things that the data set refers to which isn't actually in the data set necessarily um so one of the benefits of this approach is this ability to annotate implicit objects but also because the annotator is filling out tables rather than say doing some kind of programming or writing rdf to do these mappings like earlier approaches then someone familiar or unfamiliar with semantic web technologies or computer science may more easily uh create these mappings so uh some of the future directions that we're working on is to support additional serializations so i should also mention that our approach um supports provenance and name drops and we um right now where we pretty much create uh nano publications from the data in the form of trig uh however there's others rdf serialization such as m3 or even simpler trailers using like turtle or just rdf xml which um our current mapping tools don't create we just create kind of this um trick file but someone might want something simpler than having nano publications for example which might be hard to trigger especially if they don't know so like one of the existing directions we're working on is supporting additional serializations we're also working on supporting more complex mapping like graph joins or table joins or like even aggregates of values so for example if you have a column with a first name and a column with the last name maybe you want to create a mapping to a full name where you aggregate those columns or if you have numerical values maybe you want to do some operation on those and then we're also working on um cementite dictionary editing capabilities so um there's various platforms that supports many data dictionaries including why is and hat attack and uh standalone application so um we want to make it easy for the user to actually edit these semantic data dictionaries which may also include things like using nlp to suggest terms that they might want to um use for their mappings um yeah um i could keep going but um but yeah um yeah so um visit tetherless world slash github.io um tetherless.github.ioslash sdd to see more documentations on semantic data experience thank you thank you all right we have that okay and then i think sasha after this hello everyone i'm neha keshan and i'm working with professor jim hanson uh so today i'm going to talk about the work how we are trying to use a computer science techniques to solve a cognitive science problem that is for graduate students so i talk in a more broad manner of what we are thinking it's building a social machine for graduate mobility that's a bigger thing that we're looking forward to but the focus of my thesis would be the backbone of this system which is the knowledge graph and how we are using that knowledge graph to integrate the siloed inaccessible data available in structured unstructured and semi-structured format on web for marginalized us graduate students in stem fields so that they can use this information either extracted to make a more informed decision and why is that important is say i am a martial student i am an indian women women of color in science so my requirements will be different from someone who is physically challenged so it has to be a bit more personalized so that it can get more information that i want and the second important factor here is so that i can connect with people which i call as a point of reference who has either gone through a similar pathway that i want that i'm taking or i want to go through say a career path that i want to choose and i'm not sure whether that's possible or not but if i have someone with from a similar background as me who has taken that path then for me it seems like okay that path has been taken and it's attainable so i can take by that path and if i can connect with that person in some manner and get more information about how a particular how the day-to-day life and activities look like for them then i get more information and i know okay whether that 70 percent of that day-to-day activity is what i like what i don't like and it gives me more information to work on and make a much more informed decision so getting all these information together in harmonized manner and creating a point of reference for people especially aspirant current grad students and degree holders is what we are looking forward and we already have uh if you want to know more about this work you can look into building a social machine for graduate mobility paper and the indo the institute demographic ontology that we published and it's available online thank you and it's award winning right yes all right sasha let's try your audio again test is it working it sounds perfect sounds good um yeah i did not sign up today for a lightning talk um because i've um i'm currently right now with neha in the stage of carving out the direction for our phd thesis and as these have not been yet approved by advisor jim handler we um i'm not sure whether i should talk about these things you can talk about whatever you want to talk about tonight this is your uh lightning talks or talks of what you're interested in oh i as i understood the lighting talk was the talk about my current research direction and as this is changing i um yeah i was assuming that the um there there's then there will be then no topic for me to talk about that's okay that's okay this is your first time anyways so um well why don't you since you're relatively new why don't you introduce yourself at least sure um hi um i'm alexander i am just joined the lab this semester and um i'm working the advisor dr handler and uh currently with uh working with neho patia together to fi covered a direction for us to uh for the thesis proposal i've prior to that i was working on the cognitive co-works lab with dr wayne gray on expertise my recent work there um included the the examining the underlying assumptions for measures comparing novices and experts when looking at when measuring expertise in complex domains such as tetris where the environment is complex and dynamic in the sense that it's changing even without the player's interaction so that if the player hesitates it has to be a deliberate choice because the environment will not be the same one second afterwards um one is looking for measures how to quantify the differences between novices and experts and how they differ and one approach is to break the complex domain down into multiple subparts and investigate or sub-task and investigate each one of these independently one such measure could be for example to compare the reaction time latency to a new tetris pieces appearing on the screen and thus infer the time needed for an expert in novice to decide where to place the zoid uh however my uh work indicates my my work that i've done there indicates that these assumptions of equal task between experts and novices do not hold and if one does compare such measures one may end up comparing apples and oranges as although the measures are the same as one and one might get results that seem fine the underlying cognitive tasks that experts and novices perform qualitatively differ all right welcome sasha thank you very much you and i should have a chat sometime about 25 years ago i built a tetris auto player for george sabanko do you guys want to say something to the greater world um not to put you on the spot just that we're really happy that you're all here and uh surviving through all the challenges of covid and um i greatly appreciate that we're working together as a community through virtual spaces and masks and all of that and it's really a privilege to work with bright young minds um you know it's why we're here uh to you know explore new ideas together that's enough and i agree with that brad one of the things we will be doing exactly the mechanisms she and i are just starting to really get our act together to figure out is starting to think a little further out about what the future of the temple this world constellation is how we want to see the research go a couple of our current big projects will be ending and you know we're looking at what may be some of the new directions we want to go in so stay tuned for a lot of listening on that and hearing what's happening and uh you know it should be an exciting time we we're hoping we can begin to reoccupy this building more than this relatively you know uh um occasional way we're doing now and uh really start rebuilding community a little bit down here and see where we can get to all righty and with that uh thank you very much everybody we want to offer staff enrique or sam you want to say anything i'm good yeah i think i'm okay as well johnny want to say something i always want to say something i don't have enough time no i just thank you everybody and um you know as we look you know we're coming to the end of the term and it's always a stressful time so stay healthy and then as we're we're looking forward even with uh you know this doom and gloom about omicron and the other variants that'll show up the other greek letters you know just stay healthy and we hope that we'll be uh seeing much more of each other in person in that in the coming term all right thank you very much thank you sasha thank you fave thank you sasha for fixing your audio yes i just switched the entire machine out yes um i just have a quick question for dr handler duke and do you have perhaps uh just at least moment on your webex tonight sometime or uh to talk okay sasha wants to know if you can talk on on your webex to him yeah sometime this evening as soon as this is over i'll go up to my office and turn it on sasha that word okay thank you so much yes thank you all right [Applause] i hope that god

Original Description

Join us for a very special TWed as the Tetherless World Constellation holds our end-of-term Graduate Research "Lightning Talks." TWed Lightning Talks are a great way for the TWC community and friends to learn of the wide range of amazing research happening in the Tetherless World, and "a good time is had by all!"
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Playlist UU4rjm_R9sgRNvv9QsgH8LDw · Tetherless World · 20 of 40

1 TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
Tetherless World
2 TWed Talk: Neha Keshan on "Stress and Machine Learning"
TWed Talk: Neha Keshan on "Stress and Machine Learning"
Tetherless World
3 TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
Tetherless World
4 TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
Tetherless World
5 Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Tetherless World
6 Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Tetherless World
7 TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
Tetherless World
8 TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
Tetherless World
9 IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
Tetherless World
10 TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
Tetherless World
11 TWed Lightning Talks Fall 2019 (11 Dec 2019)
TWed Lightning Talks Fall 2019 (11 Dec 2019)
Tetherless World
12 TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
Tetherless World
13 TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
Tetherless World
14 TWed Lightning Talks Spring 2020 (29 Apr 2020)
TWed Lightning Talks Spring 2020 (29 Apr 2020)
Tetherless World
15 TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
Tetherless World
16 TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
Tetherless World
17 TWed Lightning Talks (Fall 2020)
TWed Lightning Talks (Fall 2020)
Tetherless World
18 TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
Tetherless World
19 TWed Lightnining Talks: Spring 2021
TWed Lightnining Talks: Spring 2021
Tetherless World
TWed Lightning Talks (Fall 2021)
TWed Lightning Talks (Fall 2021)
Tetherless World
21 TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
Tetherless World
22 TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
Tetherless World
23 TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
Tetherless World
24 TWed Lighting Talks Spring 2023
TWed Lighting Talks Spring 2023
Tetherless World
25 TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
Tetherless World
26 FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
Tetherless World
27 FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
Tetherless World
28 TWed Lightning Talks Spring 2024 (14 Feb 2024)
TWed Lightning Talks Spring 2024 (14 Feb 2024)
Tetherless World
29 FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
Tetherless World
30 Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Tetherless World
31 Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Tetherless World
32 TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
Tetherless World
33 TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
Tetherless World
34 TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
Tetherless World
35 TWed Talk: Thilanka Munasinghe (26 Mar 2025)
TWed Talk: Thilanka Munasinghe (26 Mar 2025)
Tetherless World
36 TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
Tetherless World
37 "Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
"Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
Tetherless World
38 TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
Tetherless World
39 TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
Tetherless World
40 TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
Tetherless World

The TWed Lightning Talks cover various research topics in NLP, including table question answering, natural language understanding, and fairness metrics. The talks utilize techniques such as retrieval augmented generation and fine-tuning, and tools like ontology, RDF, and semantic data dictionaries. The goal is to make it easy for users to edit semantic data dictionaries and access information.

Key Takeaways
  1. Build an end-to-end system for retrieving tables and answering questions
  2. Use semantic parsing to convert natural language into machine-understandable format
  3. Apply fairness metrics to machine learning models
  4. Create a semantic data dictionary to annotate and transform tabular data
  5. Use RAG to integrate siloed inaccessible data
💡 The connection between natural language understanding and other areas in NLP involves using intermediate representations to solve tasks

Related Reads

Up next
Welcome to the Next Temperamental Era
Charles Schwab
Watch →