TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026

Tetherless World · Beginner ·🧠 Large Language Models ·4mo ago

Key Takeaways

This video teaches about constructing Discourse-Aware Scholarly Knowledge Graphs for the LLM Era

Full Transcript

consider it to be there's a light. >> All right, welcome everybody to this um nth uh tw of the spring 2026 season. Um thank you very much for volunteering to do this talk today. Um uh we are not having a probably not having a couple weeks. Next week is uh the GM week Wednesday. We try not to schedule anything for that. And then the following week, I don't know. We we may or may not I'm not going to be around in the UK. Um but we probably only have one or two additional tweets before the end of the the spring term. With that, take it away. Oh, one logistical question. Are you welcoming questions during your talk or would you like to barrel on through? >> Yeah, let's barrel on through. >> Okay. Yeah. All right. So, thank you very much. >> Yeah. Yeah. Thank you. So, thank you all for joining. So, I'm so basically this is a research that I've been working for roughly five months. So, the research is basically these pseudo disco knowledge graph >> basically uh to started with what the talk about. Uh so the talk as you imagine primarily based on scientific knowledge graph construction and and the knowledge graph construction but uh in this research problem solution specifically we I had to went into uh pretty nuanced uh linguistic theories basically discuss representation theory structure theory and indirect argumentation modeling and finally uh speech act theory. uh I will like roughly mention how this been involved and what why the why did I choose this uh and before going further into the whole presentation what would be the expect outcome so in general this research would introduce an fully automatic discourse aware scholarly knowledge graph modeling I don't want to say that this would be the first one but I haven't seen any uh like knowledge graph construction uction mechanism with discourse which is able to run on fully automatically and which I will explain why is it happening and everything in later slides and second for the neurosymbolic AI fox uh this paper would introduce birectional multi-step uh neurosymbolic KGC pipeline and finally for the LM reasoning paper we are missing actually the most important one Shrihari uh did the like this proposing like representation modeling introduces an interesting idea of semantic reification inside knowledge graphs. Uh so finally in the preview so the research will introduce six or so uh like neat tools and uh introductions. First one pseudotology cement scholarly upper uh discourse ontology and ICO evocationary commitment act ontology uh sudo KGC pipeline sudo KGC MCP serance service which hasn't like conceptualized or implemented yet planning to and uh so two other like kind of neat products or like tools that I built over the research one to visualized a kg instance instances either with or without ontology itself and view the ontology isolated or with kg instances just upload internal file JSON LD RDF. uh second one is K sectoration tool which is highly specialized for the proposing ontology but can be generalized I think but like it's cool like kind of introduce and it's like a hack that you can update on a PDF without extracting anything and rerender the same thing but it's just a hack so so we've seen the preview of the uh presentation so in the later in the later slides or next slides I will go through the background of the research how discourse square discourse representation is needed and knowledge graph construction evaluation final the future work and conclusion thanks uh so first of all we need to understand what scholarly knowledge graph means and its objective primarily scholar knowledge graph are are there to connect or compose the publication metadata such such as venue, authors or the ship institute and so on with the uh concepts of the information that introduces or uses within the paper which basically you can see in this like red dot as a concept which is the most popular way uh so far that large scale knowledge graph been using uh to represent the internal knowledge and the second third and for most important part is connecting the paper that we are interested to the like broader scale scholarly scaffold with the discourse representation. So the ideal version would be of this as scaffold is the the structure would be able to explicitly represent the networking of idea in terms of relay like the evolution and interaction between literature because like it's how we like perceive papers and understand over the like course of reading many many papers. So going a little bit about evolution of SKG is basically borrowed from my RQV. Uh so we can see scholarship have been there for more than two and a half decades and the research trend being highly influenced by the predominant uh like the or the trending research structures like uh semantic web paper or the Stanford NLP toolkit. But the interesting part is even though it's been almost 5 years to this day the PPD3 kind of the first large scale uh language model introduction but we haven't seen much LLM oriented knowledge school knowledge graph construction. So this is what we are in a way trying to solve and also this like evolution is kind of interesting because it is divergent evolved over three tracks to just to give information. First track like mainly focused on metadata primarily and the citation network like middle track contains metadata but mainly introduces step the idea of concept modeling into the uh like metadata structure and finally uh final one is the scientific knowledge graph construction that I will get into that details in later slide and which is the key in this whole research. So when we talk about large scale knowledge graph, I believe most of you are familiar with these kind of products uh that you probably use daily or have heard about. So scale knowledge graph such as uh Google scholar semantic scholar A min is kind of niche one but is pretty good application and open Alex is the like rehash of the mag uh Microsoft graph that's been there till 2021 and later niche version of knowledge graph is CSKG uh but since these are primarily built on metadata uh network and citation network uh like representations. These large scale knowledge graph are primarily used on two task. First one the researchers or the institution profile discovery and the second one is this is very important keyword based cost publication discovery. Because of this reason to search uh on these kind of large scale knowledge graph for a specific paper or a specific area you have to know the keyword related to that area you just can't search on the what lexical free or scenario based searching but with the introduction of LLM now that lexical have been almost disordered uh because now you can search without uh area specifications or purely based on keyword chain and I'll explain why this being like intuitively uh made it capable by LMS and uh the second part is because of the same reason now like we can search with nuance detail and practical and scenario based setup rather than just keyword based search. So in general uh to so far like this mean this problem being solved in twofold by LLM before MCP and tool calling like LLM primarily like used their trained frozen knowledge and with the introduction of tool calling and popularization of it now like this LLM based uh knowledge graph uh interfaces are using SKG APIs to retrieve information and then derive answers. So before going into the this SKG uh a API calling we have to look into this like what happened earlier because it's kind we should kind of uh appreciate it because like LLM hasn't like even at that time even LLM is able to like derive some sort of like holistic view of area or the semantic relationship between papers. LM hasn't seen all papers together at any time of training. See LLM only has seen one paper at a time because of this whole processing happens on common set of waiting like LLM internally in like model this implicit relations. So as a human this is almost the same way that we are doing but the problem is LLM is as good as what we what it has been evaluated. So this is the main reason LLMs are good at some uh times and they are highly unfaithful at some uh promoting. So but it's just like this very interesting idea that I am proving to build the explicit representation of the semantic relations or the discover. So the second part is the how modern LLMs work basically using the APIs. So we already like defined LLMs are excelling at the in context information processing and it's increasing day by day but with the tool calling now the performance bottleneck draw down to one how good LLMs are selecting and calling tools assuming it's okay major problem is SKG API tools Because like as I mentioned earlier these SKG API tools are pre-LM they haven't updated or specifically modified into LLM support. What do I mention by that? Basically as I mentioned LLMs are good at first understand conceptual representation internally in a paper and also build the implicit discourse effort. So to harness to make LLM to harness that in the toolk itself should be capable of representing one concept modeling two most importantly explicit discord scaffolding. So we are mainly interested in the second part. Why? As you would see uh most of the existing SKG uh types such as scholar knowledge graph, academic knowledge graphs are already like representing the concept structure with the metadata but are not explicitly deriving or representing the discourse scaffold but there's a type of uh scholar knowledge graph called scientific knowledge graph. They derive this like disco scaffolding as you can see here like like this is very famous paper su deres like given a paper what is approach what are the valuations how did it implemented and so and so forth but the main problem with these scientific knowledge graphs are they are not large scale knowledge graph to use at like automated manner or use be under the LLMs. So now we have to address why is it not like scalable mainly this is like draw down into the like knowledge representation modeling issue as I understand. So so this is the like this like pink uh track that I earlier mentioned that we are interested in primarily in this paper. So like I'm like uh talking about two aspects again before geni or even with or without geni the core problem as I mentioned is the modeling of scientific knowledge as you would see in the sensor ontology itself the discourse or the the concept representation is highly derived for example this no like single entity or statement itself defines the approach. You have to connect multiple statements to build the whole approach. Likewise, evaluation and implementation and also the discourse structure is very abstract because as you again see the ontology is pretty abstract structure and also this is highly domain dependent and finally the proposing the like knowledge representation requires global semantic inference. Let me explain. So like in a paper this like approach and evaluation related state statements like scattered across multiple paragraphs and even multiple pages to to build accurate and complete approaches like whatever system that you're using has to look over large domain or large context or it should be able to do global semantic inference. So that's what I mentioned. So just imagine now we have geni. So it is some some extent capable of doing everything because ultimately they are generative. So but the main problem is with these augmented information generation it's very hard to guarantee the groundness which basically means faithfulness to the original text. Why this is important because ultimately we are planning to use this knowledge graph as a supporting or the source information for the L. So if we are making that source information is already generated by generative AI. So this like information corruption is cascading. So therefore we have to guarantee or we have to have some sort of guarantee this like generating knowledge graph is grounded or faithfulness faithful to the original text and the provenence also it's kind of simple because as long as we are using LLMs to track provenence we have to do full text search over the all entities that LLM generates that's not efficient. Next we go into efficiency. either like we can put all two like at one and get the whole knowledge graph then that would probably cause low completeness but if you're going with very like control pipeline man LLM because LLMs are very power hungry computation hungry and time costy so the system would won't be efficient enough to work as a MCP server it would be okay to work in a passive knowledge graph but again like ultimately like papers are producing like very rapidly. So if we are planning to build a static large knowledge graph that's not a good approach. So what's solution? First one, we are planning to trying to introduce new scientific knowledge graph modeling ontology approach that captures concept structure also model the discourse in a propositional structure without text augmentation or derivation at all. The second part is because of this like non- derivative forms of like scholar discourse representation now we would be able to build a fully automatic knowledge graph which is efficient almost near zero augmented near zero comes basically because of we have to anyway inference the classes and build the canonicalized relation so it kind of augmentation it's kind of augmentation and the tracking provenance finally striving for the completeness either is basically defined as high precision plus recall. Uh so before going into the detail of onlogy modeling and the knowledge graph construction, let's view the like the conceptual or the information difference between just pure concept modeling and the approach that we are proposing. Okay. So here you are seeing a concept graph that is built on very famous attention all unique paper. In a way this kind of to some extent capture the like core idea. I mean if you know the abstract this makes sense because if not it kind of doesn't make sense. Attention mechanism used for encodes that doesn't make sense. But in the statement they say you can use attention mechanism to connect between encoders and decoders. Again you see the information loss but in a way it connects the concept. But now let's look into the proposing full discourse structure. So this is what we are planning and what is the difference between this pure concept graph and the discourse aware knowledge representation. So starting from uh the first let's say the statement it should be up above somewhere. Yeah. So this is the first node the elaboration. So as you can see the statement that's elaborative statement uh bit among uh how convolutional neural networks encoders and decoders are intrinsically connected to sequence transaction. That's why like these like elaborative proposition connect with these uh concept nodes with elaboration uh relation and then we go into the approach. Let me show the Yeah. Yeah. So nice. So like this is the core they define the approach as of like introducing the transformer uh what yeah motivation or mentioning the recurrence and convolution architectures and which has evidence as you can see. So and then when we go into the evidence now this evidence connected with claims how the evidence being generated is represented by the connected concept structures such as uh it is evidence is basically running on the task of English to German translation on the matrix of blue and the material of WMT 2024 data set. So full disclosure this whole knowledge graph or the thing that you are seeing is fully automated generated automatically generated. So this is what we are striving for finally actually strived for at this point. So coming back into the presentation. So now we now it's kind of apparent like how much of a information difference happen between fully discourse represented knowledge graph scientific knowledge graph and just purely consume knowledge graphs. So then like now it's good time to understand how to model this. The idea behind is kind of simple. Basically I I try to represent how human perceived and build the global scholarly scaffold mentally while reading the paper. So basically again uh what we do when we are reading paper at a statement level we try to first classify between is this like core argument related or descriptive of uh whatever method or data set or something like definition of definition elaboration it could be some anything. So like when we understand the internal like context of a paper and when we read some paper that like also discusses same kind of entity for example attention mechanism. Now we connect these papers using canonicalized concept representation in that scenario attention mechanism. So this is the whole thing. So basically we have concept nodes we have indirect argumentation modeling nodes and concept description nodes. So that's all but actually like to build an ontology or basically we are trying to build an upper case ontology we have to formalize the uh what indirect argumentation means how what descriptive model node means and how what concept node or concept representation and the relationships build to do that I had to go for a linguistics pretty deep and pretty all linguistics. So first of all like the the concept that I just derived how human as us understand like the whole scholarly scaffold is already been theoretical defined in the concepts of conceptual semantics basically how humans express their understanding of the world and through the linguistic utterances. So there are two parts linguistic meaning of a terms of terms and the speaker intent on the on the utterances. So basically this whole thing derived in a two orthogonal spaces. First one considering the x-axis it is information structure y-axis is the propositional structure. So as I mentioned knowledge graphs pretty good at developing the information structure as a concept graph. The two missing pieces are how to build processing structure and how to connect concept graph with the processing structure. So this is what we are trying to do. First of all we doing a tad bit of different or the update augumentation to the concept graph because most of the concept like the concept graph or knowledge graph that I mentioned are domain specific. To make domain agnostic we are going with uh con classes that are domain agnostic and the lexical relations rather than the domain specific semantic relations. Luckily, there's a whole area of uh domain agnostic information extraction called CIE and there's a very good and very famous uh data set called C ERC which defines the taxonomy for the concept and concept concept representations. As you can see here the relations are lexical. They are not semantic. Because of that the the whole concept representation and concept concept relational representation is domain agnostic right which is one tick to making the pseudo as uppercase on second part is propositional discourse structure representation that's the actual novelty of this paper the work so which internally contains three parts one how to classify or define the types of propositions. Second, relations between propositions which basically define the proposition and discourse structure. And the third one is the second key part relationship between propositions and the concept nodes which actually connects the two orthogonal axises ultimately building the conceptual semantic representation into a scholar knowledge graphs. So uh yeah I'm kind of going little bit faster on these slides because they are very dry on the theory but again uh so to first model the propositional structure and classify proposition structure there's a linguistic theory called eleutionary act and eleutionary force which deres linguistic markers such as eleutionary force degree of strength sincerity which basically means which basically developed as commitment act and allocation point that we can use to guide LLM to properly classify the intent or the discourse of F statement. The second part is uh high level classification because as we seen like now we have not just statement but we have argumentation statement versus descriptive statements. To model this there's like another very famous uh linguistic theory called rhetorical structure theory which basically defines uh three concept how rhetoric basically the uh elementary discourse units for example clauses or statements are related to each other or how they are causally constructed and finally as whole schema. Okay. But most importantly for this research they introduces the core idea for separation between uh argumentation relations argumentation statement versus descriptive statements because uh in the rhetoric rhetoric relation representation of that rhetorical structure theory they derive two types of relations. One subject matter basically descriptive relations representational relations defines the preferential relations which basically derived into our required forms descriptive propositions and argumentative propos propositions. So to show what descriptive propositions what kind of descriptive propositions would this RST build or proposes as you can see here uh types are elaboration circumstances uh solution hood and etc. As I remember there are 11 types and because of this you see the same RSD theory now we have relationship between discourse propositions and the related concept nodes because that's primarily like proposition describes the core concept or in RSD satellite describes the nucleus of the uh what rhetorical structure But the second part is how to model argumentation. So assuming or representing a scientific publication as an scientific experiment. Now we have to model scientific experiment in an indirect argumentation modeling technique which actually introduced by Kleman way early I think like it's 1970 something. uh and uh I'm actually augmenting the toman's argumentation modeling using swan IOC IC SIO uh ontology and sim to build the following representation. Basically we have an idea or we have ideas to solve or to respond to issues and the ideas and issues primarily uh derived an approach. So approach uses to build evidence which proves final and ultimate scholarly claim and which has uh few like qualifiers or like the extra features such as quantifiers basically which defines in what setting that claim is true and the rebuttal and backing and warrants. So basically the warrant is the one that connects the evidence to the claim using the uh context semantics or the context knowledge. So ultimately the discourse knowledge our proposing pseudo scholarly upper discourse ontology would look like this. As you could imagine this is a multi-ter knowledge modeling. At the top tier we have argumentation nodes. In the middle tier we have artifact node and the lowest tier we have description or descriptors. And uh as you can see the relations between those tiers and inside each tier. So to look into what how these been actual realized in our extracted uh knowledge graph let's uh filter So uh so let's first look at the argumentation modeling. So argumentation should have argumentations you know claim evidence and approach. We go. So as you would see here uh approach yeah approach build evidence and evidence proves some claims uh and here the key missing parts are ideas and issues. It's not that common. Every paper has of uh complete argumentation modeling. As long as there are connected argumentation nodes, it's enough to build uh the argumentation representation. And also this is interesting part because like using the aspect of argumentation modeling alone, we can classify papers. For example, survey papers won't have any argument because it's just descriptive not. And likewise for the uh what uh idea papers or positioning papers it would only have issues and ideas or probably approaches. It won't have the claims, evidence and other parts. So it's kind of interesting but also like as you would see like it's not complete always. So the second part is I don't think we have oh we have a descriptor node. So this crypto node where is that I don't know. Uh so other relation main relation type that we modeled is the connection between uh proposition nodes and the concept nodes. As you can see here uh the proposing model derived the RST or the structure theory mentioned relations and they are like uh evidence being enabled by evidence generation being enabled by the GPU and also uh that's been contextualized in the scenario of the task. So uh like broader scheme of things this is how the proposed pseudo ontology it's been realized or populated into a knowledge graph. So getting back into the slide deck again. Uh so now we need to learn first part first of all why pseudo proposing knowledge graph is able to solve the affformentioned issues. First one uh first issue that we seen is like the existing S kgs are requiring derived information as you would like learn just few like minutes before proposing uh knowledge representation model does not require any derivation to build the propositional structure because it primarily work on sentence levels or subsent level called finite clauses. uh and also the it will define or the derive the discourse structure without using abstraction because now we are working on linguistics introduced by lexical structures introduced by couple of linguistic theories and uh it doesn't need any global uh semantics for the discourse representation because again the proposing or the using linguistic theories are running on local structure primarily. So we can just rely on the local structure but again like ultimately the argumentation like structure would be scattered but after all we have the argumentation model and finally and most importantly proposing the pseudo uh knowledge graph or the pseudo ontology make it a enables fully automatic knowledge graph construction. So next question is why? First of all, in very very simple terms, the proposing knowledge representation of modeling like schema is pretty simple because the lexical units are noun phrases and finite clauses. The propositional uh classification is primarily driven by linguistic markers that we are getting from allocation react theory, tutorial structure theory and all the other parts and proposition concept relations also derived from the discourse structure that we are building. And finally uh we can uh even do like what propositional gating before doing some semantic operation to like what control the operating operational time rather than going for n square that I will mention in very soon. uh so and going to the one key introduction that I mentioned at the prefix. So we see in the proposing pseudo ontology uh connects propositional nodes into concept nodes. But if we look at that relation in a different perspective basically concept nodes are nodes and propositional nodes are also connecting concept nodes into a single propositional node. Just imagine now that proposition node is the reified relation between all the concept nodes under it. So if you transform that proposition node into an like lat structure basically we build the embedding. Now we can do semantic reifications on that concept nodes or like we can then ext like extend that into semantic level reasoning across concept nodes and sky is the limit I guess. uh so this is the concept that I mentioned earlier. So now we have to go into the knowledge graph construction as I mentioned earlier we still have four requirements to satisfy one completeness second roundness third provenance modeling and finally the efficiency so first of all what we can't we can't just rely on PO classical NLP or non-generative models as like so apparent that that model won't be complete It's so that's been already done by some other research also researcher and but like if you go with the fully LLM or PLM based approach like this whole area of ontology guided knowledge graph extraction it is complete or it's like reasonably complete than P like NLP technique but it is not grounded or provenance being derived and it's almost always none efficient except you are doing oneshot triplet extraction. When I say oneshot triplet extraction, there's like like very uh prominent research sub area under knowledge on guide and knowledge graph extraction. Basically you put a statement and get a triplet. It's just like not the way how human build knowledge graphs. So unlike that because now we can we can't just purely work on like traditional pure NLP or pure LLM. The solution is to combine both in the on the concept or the idea of high precision semantic refinement on high recall syntactic candidate set. So high recall syntactic candidate set would generate it pretty fast because it's syntactic after all and the refinement can be done on semantic level to increase the precision. So which basically means we are proposing multi-stage birectional neuro symbolic pipeline for the knowledge graph construction. Uh so the whole conceptual modeling of the proposing knowledge graph the pipeline is pretty simple. I tried I wanted to mimic how human would annotate the knowledge graph. As a human I would try to first identify the candidate entities and try to check that those candidate entities fit into the like the given taxonomy or the ontology tet. So if it fits now we have two other problems. One we have to uh get the existing data properties for the that fitted uh entity. So I I'm I'm not assuming uh the completeness shape based completeness but we can assume it. And finally when that node being uh popularized we have to check for the inter node relations. So the beautiful thing happening in here is as you imagine this candidate extraction and the data property extraction depends on the context data. It doesn't need to have any any inference or the generative part but the classification of uh candidate identity or the canonicalized relation building need to have semantic inferencing because that classification happens on it it it has to introduce new information. It's not just an extraction. So basically now we see two sides. One lexical extractions. Two semantic canonicalizations. So this these are the two parts of syntactic and semantic refinements that I'm trying to derive. So basically like there are four stages. One I mean high level two stages like candidate structure building and relation structure building. So each candidate structure building we have syntactic high recall candidate extraction and the semantic high precision refinement followed by syntactic high recall discourse structure building. So which basically means relationship building and end with semantic high precision but like actual pipeline is little bit complicated than that. So let's go with an uh running example to understand what completely happens in here. Uh so I'm explicitly not trying not to go get into algorithmic detail. Uh just to show what kind of augmentation happens over the pipeline. So first sentence augmentation from the like given paragraph in here. Again I'm sticking into uh attention all you need abstract as we seen the demo and center segmentation module would separate out the sentences again using uh what lexical rules and then we separate sentences into propositional structures. Basically here proposition means finite clauses in like simple terms uh like clause that contains subject predicate and object or subject verb object. Then we do open information extraction as you would see here like in red colored one other introducing new information in the pipeline. uh we extract or the define the uh like spans of named entities. So this is purely uh syntactic operation basically try to derive uh adjacent consecutive uh noun phrases or proper noun phrases in the text. And third part is that little bit complicated in algorithmic sense but it's pretty easy to understand local correerence resolution because like in a like when we see noun phrases there are two parts some noun phrases are uh what grounded noun phrases that can stand alone but other noun phrases need to have referential unit to and grounded information that's what basically what core reference resolution try to do. So as you can see here like we have grounded uh in this example we only have oh at the very bottom we have a core reference resoluted one uh yeah so then we go into the semantic candidate refinement where we at one shot do three parts actually like this is like highly pipelined uh DSPY base uh LLM refinement but uh basically we do referential named entity refinement, grounded name named entity refinement and finally named entity classification. As you would see the result would be like that. We would get classes basically these are like semantically uh semantic representation of the rack named entity and also this like red color complex uh recurrent is and drop. As you could imagine here, we had that complex recurrent as a grounding node and this uh but grounding node refinement would detect. Okay, that's not an uh very good grounded node because if it it was like recurren neural network yes that makes sense but it's just complex recurrent. And finally uh this is that a linguistic marker generation is IOC attribute ICO eleutionary commitment act attribute prediction for each proposition we predict what is the force what is the strength to talk level polarity and the modality of the proposition because at the later stage these linguistic markers work for classification as you can see Here this propositional discourse type classification primarily work on the like inferred proposition markers to detect this is an argument and which is a claim. Uh then uh we again get into the second phase of the knowledge graph construction the discourse representation. The first part is uh argument to artifact canonicalization. Again this is semantically required process because of that we are using LLMs. Uh and you can see the relations that LLM generated uh defining what like the named entities being contextualized in the proposition. uh then uh we would derive artifact artifact uh relations as you would see in here again LLM based uh prediction based control prediction and uh then we again move into a syntactic uh argument discourse gating. So the whole idea is this okay like as I mentioned there are three types of relations artifact artifact descriptors to artifact argument sorry there are four argument to artifact and the discourse structure only builds on argument to argument so if we haven't done any gating we have to check the like n² combination space to check which like arguments are like visosely connected to each other. So without going into this n² operation we can do gating on two levels. What connectiv one connectivity gating which basically means one causality based connections basically you have to have an approach before evidence because you will generate evidence based on approach. So like that kind of causality I use for the connectivity gating. And the second one is predicate which basically means domain range relation constraint based uh filtering because like uh what uh proves connects evidence to the claim none other and then based on that I'll do uh refinement on the argument discourse but this is not LLM here I'm using NLI models for the ent prediction. rather than just classification because ultimately it's just like given two statements we should be like the if it is if they have an like expected uh relation we can set the relation sorry proposition into the hypothesis and u put the relation into the premise and ask would those uh two propositions entail the premise and that's how we get the argument argument discourse and then uh we do an like kind of additional step of refinement on arguments as I mentioned argument nodes does not stay alone because if you have just have the proposition or the approach it doesn't mean anything because we can have an approach as in a description or the elaboration to have an argumentation you have to have the connected structure. So therefore we separate or we we redact uh or all the like isolated argumentation model argumentation propositions into an descriptors and try to predict the descriptor type. Uh yeah that's what happens in here and that's why we like make and that former claim argument into an descriptor of elaboration. Uh so all in all this is how knowledge graph construction happens and this is the exact pipeline that I used to generate that knowledge graph that you see. Finally for the evaluation again I'm like fixing into the four requirements that I mentioned again and again completing against bag of cont bag of concept performance which is pretty famous and also pretty bad that I will show and the node classification accuracy again for the completeness for the groundedness we use uh node level hallucination detection that introduced with in the paper of uh text by professor or Dr. uh provenance and data quality validation using shackle or more specific structure of shocks uh to do multi-axis data quality validation and finally the efficiency just runtime analysis I haven't done it yet uh so one part so all these data that you are seeing or the performance result that you are seeing primarily on a very small basically five paper annotated calibration set that I'm I'm still calibrating the whole uh pipeline Uh so first one so like this is a bag of concept performance I defined three forms basically uh leanstein text similarity uh with fixed and dynamic thresholds as you can see here like this is the core problem with bag of concept it's just set intersection or union at performance or the scoring so if you can set a dynamic threshold with the partial matching you clearly get a one but it is not the correct one. So because of that and also like because of the set matching that doesn't guarantee one to one matching. So in here I have to derived like kind of more stable form of a bag of concept performance using Hungarian bipartite matching algorithm which kind of inspired by an paper called DRT which is like image classification paper or object detection paper. uh and as you can see here like which is pretty like stable and also I'm not just using cosine similarity I'm using node reranking model to derive the relation between or the similarity between two entities so the current performance seems to be good but I'm not that satisfied this should be way better uh because like ultimately we are trying to strive for the full completus uh specifically that you can see in the like node classification performance. It's like not that good uh specifically in the proposition structure because like it's kind of complicated and also like my annotation was pretty bad at that time. So I was like sleepy uh and but the interesting part is the grounded groundedness and the data quality validation. So for the groundedness I did the hallucination evaluation. So as you could see in even with this like smaller paper account we are not getting any hallucinations. I mean practically there is no way but like anyway like we are evaluating it. Uh but uh for the data quality measures I'm evaluating again four axises. one completest which basically means the uh expected discourse representations and the uh descriptors to concept representations. Basically if you have an elaboration that should have elaborate some concept like that's why we are getting low score for the completeness because some description like propositions are not connected with concepts and uh literal validation which basically means English should be English from the literal and provenance because like we are defining the provenance as of from where uh the like the note being derived from exact content text and what paper pro was generated and uh pro agent to derive what version of the pipeline that's been used and also finally the range validation which basically validate the domain and range of uh relations and but again this is not the actual validation that I'm planning for the final version so which contains three parts one gold standard data set which contains which would I derive from AMSR data set which is a uh what open reviewer data set uh which has AKB AKBC conference with 27 papers which is kind of interesting because AK BC is academic knowledge based construction uh oh sorry automatic knowledge based constru construction uh and I'm planning to evaluate on I I'm planning to annotate all these 27 papers manually and do run all those four axis like requirement axis evaluation but like to like evaluate on large scale uh I'm trying to or planning to introduce an standard data set with like distance supervision basically this is like a recent evaluation method introduced called mine which evaluates the induced uh information completeness so uh I'm trying to evaluate on that and finally ablation study uh on the goal standard data set to show uh whatever the like the techniques under each step is the optimal hopefully I have enough time to do that but yeah so the future work again as I mentioned I had to build the evaluation sets to some extent I have extracted most of it uh I have to run the fulls and and most importantly so far you've seen this whole knowledge graph construction ction within an paragraph but to do that over a paper we have to have global cop resolution that I have to implement I mean that I had implemented but I haven't connected with the pipeline yet uh but also we are trying to introduce another paper from this research as of like resource paper which ultimately build this into a like actual knowledge graph which connects across papers and canonicalize the concept nodes basically global economic class using like what CSO classifier and also implement that as a potentially a MCP serve and service uh and also work as a post hawk uh like data representation to the semantic scholar API for the paper metadata and broad scheme of thing we are also working on the knowledge discovery on the same like sudo kg which kind of we are on the path of like deriving how research been evolved and minimal contribution unit of a given paper and also the concept semantic drift over the paper over the research directions actually and finally this is an interesting idea that I came up came up over the like just come into my head why not do a latent knowledge retriever because like like right now we are just doing like we are embedding the like label or RBF label value of a node and check the cosine simulator but which basically means we are losing a lot of information data properties and the concept uh or the basically class ontology information and lot of stuff. So but if since we are we have this like semantic reification we could probably do latent knowledge retriever I don't know I just like an idea and for the conclusion uh we try to introduce fully automatic discos knowledge graph I think we've been successful at that and as you seen proposing knowledge graph is information information rich compared to the previous concept graph apps and find then like with the in in the middle of the research we somehow came across this idea of semantic reification and finally like the research introduces a nie KDC pipeline for the zero or like uh almost near zero information augmented knowledge graph construction and that's it thank you Any questions? >> Where to begin? >> How many papers in the evaluation set? >> Oh, five papers. Yeah. Asking about the current evation, right? Yeah. >> It's five papers. >> Yeah, that was the left hand column of that. >> Oh, who missed that? >> I was looking for that when he hit that. >> My focus on more topics. Although I was explicitly trying to make it like different as possible. So one paper on knowledge graph evaluation as a survey paper, one propos like position paper, two full research papers and one short paper. What topics? >> So they are mostly on knowledge graphs >> construction >> kind of yes. for one transformer paper. >> If you threw some medical research paper, an engineering paper and would that still work? >> We hope because the concept structure is generic enough. It doesn't contain domain information. But again like it kind of depends on how good the semantic refinement is. Anyway, we would get the high recall like entity set because in any domain like this concept node represent as a proper noun or proper noun segment. >> Yeah, I think I kind of understand what you're doing but I don't quite get this discourse piece. So you mentioned dialogue acts but I didn't see any in your examples any >> oh so >> anything out of that I mean somebody disagrees with with like what you claim or or questions claim or >> I'm not using dialogue to if I'm if I'm using dialog that happens I'm using speech which like single indirect argumentation without any like call back like in dialogue basically because within the paper they try to build an indirect indirect argumentation means first define the problem and then give the claim right but but but what I'm guessing that one paper may argue that another paper is wrong about something right and so very often that is you don't say you know the guy so and so is wrong I am right I mean that's that's probably people don't say that right >> yeah yeah >> they would say things like you know so and so this way but this only works in this case and and I kind of agree with with with what they say however actually the I have a better solution that is a an argument against something. So it could be quite subtle, right? And you have to get that and somehow have to have some kind of a weight to it. Is it complete rejection, partial rejection? >> Oh yeah. >> Yeah. back. So this it's it's really like conversation except it's in the on long distance and indirect. >> Yeah. >> People are not talking to one another but talking sort of indirectly talking to some global audience by saying this guy is wrong. Okay. By by being indirect. >> Yeah. >> Right. And and I think I I thought that that would be most interesting piece rather than just extracting ontology right out of actually I don't know what you're extracting. Is it an ontology or is it it's it's an argument structure that you extract? >> I'm extracting argument structure with the like connected concept no structure. like I totally agree with you like if I'm only deriving the argumentation only but like this like descriptor like nodes explicitly is there to like dull down this like like global prito argumentation because after all like as I mentioned the ideal uh like structure is how like ideas evolved. So to do that at least in this research like we would imagine if we have the initial idea and the approach and the evidence and the claims and the like the couple of like tidbits like quantifiers and warrants would be enough but also yes paper also defines the like how other papers hasn't been able to do that or like in very indirect manner or like very cautious Okay. So that kind of uh statements goes under descriptors like elaborations, contextualizations, causes and so on and so forth. We preserve it but we are not explicitly using that into the structure. Okay. Um, so I think it was somewhere around 550 you start doing that that you start incrementally showing how you're building out. >> Oh yeah. >> That that that kind of in injecting that that the semantics. Do you have a earlier when you were showing the graphs that the visualization in your idea star tool do you have a uh like a visualization of that? So go to go to around 50 or so when you start building I think it was around 50 >> like here. >> Well yeah you you you start in you basically you do your initial parsing and then you start building. >> Okay. So, a couple things. One, I think to your uh to your viewer, I'm pretending I'm I'm last semester. Okay. Um I think it's it's really hard for us to see your red highlighting. >> Yeah, I just feel it. >> You know, if you did something like bold or whatever to kind of show how you're incrementally building out, that's just a that would help because it was really hard to kind of pick that out. I saw that you were doing it, but I crossed my eye. >> I'm impressed that you can see that. I cannot tell. Like, knowing that it's there, I'm like, is it though? >> I'm I'm sitting I'm sitting 42%. >> Oh, wow. I guess that's it. Test out your It's super visible on your screen. >> This is rule rule. It's It's rule 3A and it's in your presentation forum. You should always check it out. But anyways, you can you can see my point how you want it. >> I didn't realize it. >> But but but it but I mean it's it's very cool how you're building out. But what I was as you were doing this, what I was imagining is that you had like an incrementally constructed visualization. >> The actual thing I'm seeing like >> Yeah. Yeah. Yeah. I would illustrate that as you're going along. Okay. Um that would that would help. Um I guess I've gone down this path so I should keep going. Um you introduce a couple you introduce in passing a very novel tool which we've never seen before. That that visualization tool. >> Oh yeah. That I really >> Yeah. That would have been worthy of an entire twin by itself. Okay. >> Um because we've had Exactly. This will be like the third that we've had in 15 years or 16 years. Okay. So, yeah. Yeah. When you got it like that's like, oh, as I'm going along in this war, I'm going to drop this bomb. See it? I'm the next thing. You know, this was this was a very innovative thing. Um, >> which we should talk about. >> Exactly. >> Not now. But it is a >> basically it's our RDF viewer, right? A very >> Yeah. >> Cool RDF viewer. Um the generic or >> it's generic you can put any nontologology on any like turtle instance. >> Yeah it's like early it's like the newest RDF >> problem with RDF visualizers. It needs to be like parameterized. >> Yeah. >> Show like specific entities or group of entities. I've I've only ever seen one that I really liked and and Jamie refused to. It's it's like it's like Apple making sure that old versions of iOS don't work anymore when they introduce sudden you buy the one when Jamie introduced why is all of a sudden she didn't support RDF viewer which is actually amazing which I'm reminded every time I give a guest lecture up the hill I I show demos from RDF viewer and say I wish I still had this anyways um this is very very cool uh do you have this in GitHub somewhere >> so yeah like it this has GitHub and already Docker build I'm planning to like the build an image and put into the like the Docker hub that others can just download and run as a container. Nothing much. >> So it' be really cool to see um this demonstrated with other examples. >> Yeah. >> How long have you had this around? I mean I had build this in the middle of I just wanted to see the turtle files because of I built it. >> Well I mean it's nothing really really good like this exists. >> No. >> Okay. You know this is this is very innate. This by itself is a could be a resource paper. >> Yeah. Um >> yeah, you would have to have that faceted angle uh or else it wouldn't it wouldn't fly. >> Yeah. And I think for large graphs it also be a problem. >> I try to make it scalable as so this is running on like what lazy rendering on D3JS with condor's uh like condor's piplet representation structure behind it. And this is like what uh what is that like? >> Thank you native. >> Yeah. So kind I think I think go ahead just just two comments. Uh one thing is I think if you know the schema in this case you know the schema because I have the ontology. Uh you can have like an ontology driven plotting. So you maybe show the entities and then you can potentially see the instances associated with the entities. So that helps. But if you don't like that, right? So you have that already. But let's imagine you have a graph and you don't have the schema. >> Yeah. >> Right. Which is comma, right? >> Like just put in the turtle file. >> Yeah. Any turtle file. A large one. >> It will still show. >> This is larger. >> Not that large. >> Yeah. But if you have like a large one, you could have like uh uh some embedding approaching in place to extract common entities and maybe cluster the instances together. >> You already been done. >> Oh, you did? Yeah, I can cluster by types >> the IRI also like it's like I have different view levels like I can show complete uh object properties annotation everything or >> yeah this is this is >> but that's the ontology view right >> no >> you have the ontology view >> oh like this ontology view means like different levels of like ontology only the class structure, class structure plus object properties, class structure, object, object properties, >> good enough, but uh for graphs that have >> Yeah. >> So, uh yeah, >> yeah, and I put it in the chat, but are you connected to any reasoner? So there was a an ancient entity uh relationship browser uh early in the description logic world and it was connected to a reasoner and then you could see the updates being added to the graph which was crazy cool. >> A little this has existed for what like two months. Come on. >> Well, but he's so fast and so energetic you know maybe he'll do it. >> Oh yeah. I don't think he seriously I don't think he's ever Yeah. You could you could try to submit a resource paper this on that. I you >> actually it would be interesting to see what kind of comments you get. I mean I think you should think about it. >> You get access to because it's like >> yeah it's I don't know if this is like of value anymore these days but you can try. >> It's definitely a demo. So >> so >> yeah. Yeah it >> you know it might take off just like Widoko did. I mean, woko was pretty simple, but uh it really kind of took over. >> Yeah. >> But yeah, Vokco was such as was a piece of excellence. Um well, yeah. So, it would be really stupid if this got overlooked because like the the semantic web world has existed for without really good drive ups. And it's like how does the semantic web do their world their work without having something really good? And um so this is cool. Um I think in the interest of time Paul quits but I did want to ask >> can you talk about what like are you submitting this to a venue? You're doing a talk a formal talk soon. >> You mean >> yeah your work not this but your work in general. Oh, so the that whole thing that I mentioned would go under as a full length paper into ISWC and the MCP the structure and the global canonical organization with the full connected knowledge graph go as a resource paper into ISWC. >> Okay. Um that's pretty impressive. Um the uh when this gets accepted um you're going to I you're going to want to uh figure out a strategy for it because this is your presentation kind of boils two oceans or three. Uh there's a lot here. >> Yeah. >> And you you you're you're gonna have to figure out a strategy and and literally it's it's it's too much for I mean you crammed it into a little over an hour really quick talking. You zip and coded it. >> Um >> you you need to figure out uh a kind of a a strategy for presenting this in a in a very concise way but not missing anything. How do you how do you what are the headlines? How do you paint the pictures? And then some of the mechanics that we talked about last fall in terms of helping your your viewer keep track of where you are. >> Oh yeah. >> Yeah. You know, you know that because they they they sort of need to be able to have these mental milestones as you go through. I'm here now. I'm here now. You know, he's explained this now. He's explained this now. So, kind of help contextualize as you go through each piece where where you're at, you know, that I think that helps the for long talks. It it helps the user or helps the viewer to you want to figure out where you are. >> Yeah, I try to get get like bid crumb like the header that like Jane >> I know, but >> but yeah, I mean and and I'm not saying your argument wasn't logical or anything like that. I'm I'm not I'm not criticizing that. I'm saying when there's a lot, you need to sort of like this is what we know so far. This is where we are in our story. Okay. Kind of bring them back. So, now we're going to talk about this. You see? >> Yeah. Yeah. >> Kind of have to take a break and or you know, kind of take a thing. The other thing to think of for this forum um for tw uh you you want to think about how you might have discussions and stuff like that. I mean that's that's it's well it helps it helps keep both but it went to we would we you'd be still getting through the first section of the paper if we actually >> Yeah. Yeah. Yeah. There's a lot here and again that's not a criticism when you got a lot to talk about. It's hard to So you did the right thing by saying don't ask questions. >> Yeah. But but part of the the the sharing here is getting other people's ideas. So, making sure to leave some time for questions is always a good strategy.

Original Description

Scholarly knowledge graphs have traditionally linked publication metadata and extracted relations to support discovery across large collections of academic literature. However, these representations primarily capture concept-level relationships and largely omit the discourse-level knowledge encoded in scholarly texts. At the same time, the emergence of large language models (LLMs) has fundamentally changed how scholarly knowledge is accessed and consumed, while the underlying knowledge infrastructures have remained largely unchanged. In this work, we introduce a new class of scholarly knowledge graphs that integrates concept-level structure with explicit representations of scholarly propositions. We model scholarly discourse as networks of interconnected propositions grounded in concept nodes, allowing propositions to function as semantically refined relations between concepts. This representation captures both the conceptual structure and argumentative structure of scientific knowledge, enabling reasoning across discrete symbolic representations and the semantic spaces leveraged by LLMs. To support this representation, we propose the Scholarly Upper Discourse Ontology (SUDO) for modeling scholarly discourse and develop a bidirectional neuro-symbolic knowledge graph construction pipeline that combines high-recall syntactic candidate extraction with high-precision semantic refinement. This approach enables the construction of discourse-aware scholarly knowledge graphs designed to support both human knowledge exploration and LLM-based scholarly reasoning.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Playlist UU4rjm_R9sgRNvv9QsgH8LDw · Tetherless World · 40 of 40

← Previous Next →
1 TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
Tetherless World
2 TWed Talk: Neha Keshan on "Stress and Machine Learning"
TWed Talk: Neha Keshan on "Stress and Machine Learning"
Tetherless World
3 TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
Tetherless World
4 TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
Tetherless World
5 Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Tetherless World
6 Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Tetherless World
7 TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
Tetherless World
8 TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
Tetherless World
9 IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
Tetherless World
10 TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
Tetherless World
11 TWed Lightning Talks Fall 2019 (11 Dec 2019)
TWed Lightning Talks Fall 2019 (11 Dec 2019)
Tetherless World
12 TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
Tetherless World
13 TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
Tetherless World
14 TWed Lightning Talks Spring 2020 (29 Apr 2020)
TWed Lightning Talks Spring 2020 (29 Apr 2020)
Tetherless World
15 TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
Tetherless World
16 TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
Tetherless World
17 TWed Lightning Talks (Fall 2020)
TWed Lightning Talks (Fall 2020)
Tetherless World
18 TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
Tetherless World
19 TWed Lightnining Talks: Spring 2021
TWed Lightnining Talks: Spring 2021
Tetherless World
20 TWed Lightning Talks (Fall 2021)
TWed Lightning Talks (Fall 2021)
Tetherless World
21 TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
Tetherless World
22 TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
Tetherless World
23 TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
Tetherless World
24 TWed Lighting Talks Spring 2023
TWed Lighting Talks Spring 2023
Tetherless World
25 TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
Tetherless World
26 FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
Tetherless World
27 FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
Tetherless World
28 TWed Lightning Talks Spring 2024 (14 Feb 2024)
TWed Lightning Talks Spring 2024 (14 Feb 2024)
Tetherless World
29 FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
Tetherless World
30 Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Tetherless World
31 Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Tetherless World
32 TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
Tetherless World
33 TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
Tetherless World
34 TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
Tetherless World
35 TWed Talk: Thilanka Munasinghe (26 Mar 2025)
TWed Talk: Thilanka Munasinghe (26 Mar 2025)
Tetherless World
36 TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
Tetherless World
37 "Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
"Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
Tetherless World
38 TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
Tetherless World
39 TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
Tetherless World
TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
Tetherless World

Related Reads

📰
Will Developers Need LLM Integration Skills in 2026 for Success?
Developers will need LLM integration skills in 2026 to stay competitive, learn how to integrate AI into your workflow
Dev.to AI
📰
I Trained a 471M-Parameter Language Model From Scratch on One RTX 4090 in 100 Hours.
Train a large language model from scratch on a single GPU in under 100 hours, leveraging recent advances in AI hardware and software
Medium · LLM
📰
Masking PII Without Losing It
Learn to mask personally identifiable information (PII) without losing its value using LLMs, ensuring user privacy and data security
Medium · LLM
📰
Build a Career in Artificial Intelligence : AI Mastery Course in Telugu
Learn how to build a career in Artificial Intelligence with a comprehensive AI mastery course in Telugu
Dev.to AI
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →