TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"

Tetherless World · Advanced ·📄 Research Papers Explained ·3y ago

Key Takeaways

Explores the limitations and possibilities of bibliometrics in studying scientific disciplines, including citation analysis and research concept relationships

Full Transcript

so good evening everybody and welcome to this snowy twed so we're virtual today um missing out on the wonderful pies and pizza pies but we'll make up for that in the in the near future um uh we're today we're very happy to have Brenda Thompson uh to know this world uh PhD candidate um talking about bibliometrics um we're all looking forward to that uh just a couple logistical reminders were recording this uh tonight uh Brenda I I think you're welcoming of questions is that right yes absolutely be a discussion sure okay so it's a discussion throughout um and just keep in mind everybody that you are being recorded and with that I'm going to kill my share okay and Brenda you can start sharing and take it away oops this actually works it does okay Irene looks great hit it I like this one much better okay all right well thank you everybody for being here first of all um thank you for the nice introduction John of course and um so basically yeah bibliometrics um here to kind of talk about the limitations and the possibilities and um so without further Ado let's kind of let's start by um okay and my keyboard just died okay maybe I can't let me see okay there we go so um start with a quick definition right so what is bibliometrics I want to make sure we kind of get all get off to the same star here and so um it it really is just when you go to a library right it's it's essentially com it's the statistics that they use on essentially a paper compared to a paper um so with that you know what is it all really actually mean to us as scientists or as researchers is that for funding reasons for our own work for to advance our own research right we all kind of get locked into having to do a literature review and so with that right you always kind of start with a with going to you know web of science or scopus or Google Google Scholar right this is kind of and these are actual bibliometric searches so um and so what we do with those uh when we're looking at a literature review is you know we get our search results and then we kind of evaluate what bibliometrics can provide is an evaluation of those sources right so that are actually from the results that are returned from our bibliometric search and then with that we move on to trying to identify themes and looking at the dates and gaps in the literature so that we can actually then of course um provide our outline and and then ultimately write our literature review um so building metrics is very strong in the first two and then by the time we get to identifying themes debates and gaps is where though bibliometrics is there and available um some of it actually this is where it starts to falter we don't really have those aspects involved so what does that actually look like right so what is happening with bibliometric search right we start with doing the actual search and what we get back is the literature and then that literature that we do have should be you know essentially our uh relevant to what we are focused on learning or trying you know trying to get get at at the base level and so what we get from bibliometrics is exactly this we get the sources the authors and the documents and what's really kind of easy to get out of that is of course you know we can find what's the most relevant who's most cited um and you know what what is happening within those papers actually right so how many papers are being published per year and we have um Bradford's law here in the sources which is kind of an interesting uh limitation that is that that exists and basically what Bradford's law says is that um it essentially estimates estimates the exponentially diminishing returns of searching for references in science gym so basically if a researcher has like five core journals um that they're interested in right five core journals for their specific science domain um and say that in one month there are like 12 available uh articles of interest in order to find another 12 articles of Interest they're going to have to go search in 10 journals 10 more journals so basically so what happens is that after you've you know after an author or after a researchers gone and looked in 5 10 20 40 journals then there's you realize there's no point and so what happens is that the library has decided you know the Bradford's laws explains that libraries uh for libraries it's sufficient to have just those five quadrons so we're that's a big limitation and then you'll notice with authors we have a lot because long and what's what's happening there is this is a law that describes the frequency of Publications by authors in a given field and so basically so for the number of authors that are publishing a certain number of Articles um this is a fixed ratio to the total number of authors publishing a single article so basically as the number of Articles published you know goes increases what happens is that authors producing Publications becomes less frequent so again limiting factors that happen as a result of uh you know bibliometrics it's just the nature of people if you will so and so now so just so that we're all clear on where we're starting from I'm kind of going to get touched on some real super Basics here and um let's see so this is uh scopus which we have access to Via either API or just because or you know RPI staff and students and um so this is scopus.org and what we have here right is you'll notice um when you go to this to start your search you have doc you can search by documents authors you know you can search the researchers and uh any you can also search by affiliation and with that then you can also do this add addition of search fields which is the large list on the left so you can really kind of narrow down your search when you first get started we're probably all very familiar with this I would assume um but and so what I did was a search on the origins of life and you know with with no constraints just for this example and so for this um I've got 6 500 results and if we're researchers doing a literature review nobody's reading 6500 documents um so to get at that further what we have here then on the left hand side is the the facets of the search right and these are kind of determined by the metadata that's available uh within you know whichever service in this case scopus that you're uh subscribed to here and so we can then of course narrow the search down greatly as we move through and refine those facets so um just for argument's sake here I oh let's assume right that um I've managed to use the facets to whittle this down to like 300 so it's something reasonable for literature review um but what we realize is over here what I've highlighted is uh we have some documents as well as these secondary documents will actually be those documents that are directly cited in the original list of documents you can also wind up looking to see if there are any patents underneath the uh works that that are presented in your final results and what happened um yeah okay and then I want to point out specifically there's this option here to analyze search results and so this gives us bibliometric search and what that kind of looks like basically is something this simple right so this is what you're looking at is documents per year by source source being which particular journals which I've got highlighted at the bottom in the Red Bracket here and so what you'll see right the source here is journals so you know what's who's publishing on uh what's the number of documents by Journal uh for origins of Life over time so that's basically what we're looking at what I want to point out here is if you'll notice we have the blue line up here and as well as the red and um so even though you know uh metadata can give us problems in in this way what I know what's happened here was that uh if we look down at the names of the the journals in particular one is INS in the biosphere and the other one ends in biospheres but around 2007 we decided there were more biospheres to explore so the journal just you know slightly changed its name but this kind of error could be easy to miss so something to think about so further on the you know analysis of search results it's going to give us like I said earlier right number of documents by author docs by affiliations and these affiliations just word to the wise you know if you are doing specific research this could also be a great way to look for um people that might or in Industries or uh employers potential employers that's what I'm trying to say and finally you know they give you the the ultimate pie chart and uh again this is just subject areas and it's going to be based on journals again as to what is uh you know based on how that particular Journal is already classified again this is all locked into the metadata so with that um still a peek at oh yes one big thing this is still a paper to paper analysis we haven't even gotten into the context at all and so just give us another quick background I'm gonna pick on Professor Schaller um and so you know when you go to Google Scholar you can pick this up you can sort by title you know you can sort by how many times this was cited um you can look at the year but ultimately their bibliometric tools are just this right so again citations how many citations do you have and we have the H index making an appearance and uh and then you know ultimately how many papers and when they're cited across the years so still no context so ultimately let's take a peek at what other bibliometric tools are available and for that we have an external program known as Vos viewer which I'm gonna you know this boss viewer requires a Java runtime environment it is a free service and does provide a lot of network analysis again this will kind of let you dig into some of the context but I can't run it on mine with that because of java issues so I haven't dealt too much with that next is site space now this is one that came out in 2016. it's uh external but it is a paid for if you want to do more than 50 articles at a time so it's it's out there it's available to to do some analysis with whereas it'll actually let you get into the context of the papers a bit but not not great and finally there's um excuse me python has a package and R has a package as well so give you an idea of what that looks like again another bibliometric tool this is ours bibliometrics program her package excuse me and the um what we have Happening Here is been off To Each corner right describes the quadrant and again you know origins of life in my particular case I am very interested in only the only origins of life on Earth so as you can see the bibliometric search that I did actually returned very little you know Earth is there but most of it has everything to do with astrobiology and space Etc I'm particularly interested in drilling down to Prebiotic chemistry and that's almost you know that's weekly developed and marginal uh in that lower left corner so I would need to certainly refine my search to get at what I need and again this still if you look at it in the upper right corner you'll see that one of them is article right so again we're still looking at a paper to paper analysis so I explained this this earlier so what I'm suggesting is that we get a place of metadata games and with a few other um options we could do this we could actually add on the mapping of the science which is what I think when we do literature reviews that's probably what we're actually looking to do um we need to get at the concepts in the literature the knowledge and the literature and kind of knowing what's happening socially will absolutely help us understand and give us more of what we need for um getting at our grant funding or things of this nature so let's kind of talk about the workflow on this um so the data acquisition starts with the viviometric analysis right this is that and then what bibliometrics is really providing at this point is these citation networks right and um the site you know we established these right citation relationships and it goes on to essentially build a a something that we can look at as the knowledge landscape if you will of of the topics uh in the academic articles and what I'm suggesting is we move to you know incorporating cement more semantic analysis and to do that we would have to start with once we've got the the data then we're going to have to first of all initially clean it and then and then we can move to doing some Network science um to work out basically a hierarchical step of how the science has progressed and that'll be a combination of network uh of network science as well as some processing in with natural language processing and then finally um to get at the evolution of a particular topic and so we can look at emerging uh research and things of that sort we could lay it out and do some evolutionary pathway so just to give you an idea so again here's the bibliometric results we're going to have to initially clean all of this data and to do that of course I'm not sure who knows what the processes are but we can you know essentially clean up all of this data make the words uh tokenize the literature itself into actual words or paragraphs depending on how we want to analyze and then it'll come up with some summary statistics before we move on to doing further analysis and at this point is also when we can really start asking some questions of the data and determining what needs to be well we could actually get out of this and so with part of this um we need to do some feature engineering right so if we are able to tag these are common NLP methods right so if we can actually um do the parts of speech tag tag a noun is a noun um then we can actually and get rid of all of those other uh is a um then we can get at what is actually happening in the paper um additionally we could move on to named entity recognition identifying you know those from the metadata Fields themselves pulling out that information that's also in the paper and finally we have to get around to something like a Word Sense disambiguation right so am I talking about um am I talking about when I say base or Bass right spelled the same way b-a-s-s am I talking about a fish or am I talking about a guitar so we have to be able to pull that from the context as well um next up we would do some topic extraction and modeling um and again we're going to look for identifying those main topics and to do that we'll have to use some algorithms before we get around to doing the actual modeling um and we have to construct a corpus and things of this sort to get through the natural language processing part of this um next up would be then after we've got you know our topics clearly extracted and we have an idea of what's going on then we can rearrange the data again to um uh to do some more Network science on that and um I'm going to start with again some sticking with some real Basics so we're all on the same page so um the blue dots here um are essentially components in the network so let's assume that each one of those round blue circles is a uh is actually a paper go with paper um and so these papers right these components are all part of one network and so as you can see so what connects the papers would be citations so one paper citing another paper um and so these components though part of the same network aren't connected and this might be because you have different topics going on or could be a Litany of other issues um but ultimately what we then would move to is kind of pulling out those communities that exist and and in this case maybe what I've circled here this could be papers that are all on the same topic um or have the same author or depending on which metadata field you are you're exploring and so um these components uh lead to the communities but what we also need to think about too is the centrality with things like this particular um node would actually be a would have a high centrality because it's connecting this other paper that would otherwise not be part of this component or this community um without you know without this particular without that particular node so um so again so if you are looking at you could be looking at this from the perspective of a paper or the perspective of an author or any of the actual metadata fields that are available so there's a the ability to drill down to a singular paper or a singular Community but there's a lot more information here so let's kind of take a look at what the centrality would mean if we can actually you know when we actually get these papers uh into this network structures and so here um we have this we have degree centrality which we can see if the gray node in the center right if if that particular paper wasn't there then then none of these other papers are connected and so and then we have a between the centrality that we can check where we can see that there might be two communities or two topics that uh are connected by this one particular author or this one particular paper meaning that without that particular contribution um you know they would be two entirely different things so as someone doing a literature review you know you might want to just go pick that particular thing to explore and again then we have closeness centrality which in this case the lower um the lower one here would um not even be connected and this could be a very a highly influential paper um that without the gray Central one wouldn't be connected again to the rest of the network finally there's eigenvector centrality which this gray node in this case in an eigenvector situation when you pull that particular centrality out of it what it's measuring is it's it's an important paper that is connected to other highly ranked nodes so other highly ranked papers or other highly ranked authors so this again this is something you someone or something you might want to take a peek at in order to get further into your own literature review so with that you know let's talk about you know we look at the all of the if all of these uh the circles are actual nodes right are papers in the network um and we're doing literature review um if we look at how this is structured what we'll see what we notice here is that we may be able to get away with just reading these two papers to determine whether or not um we need to to further Trace out this this network of papers and additionally I mean the worst case scenario we could read all four of these and decide whether or not we need to delve further into these other cited papers and so that's basically what I wanted to present on I wanted to talk some more about um you know metadata Fields here um one of the ways you know how are we accounting for in bibliometrics how are we accounting for things like semantic shifts um and for example um I'm a dog lover so I'll go with you know it used to be way back when we all dogs were referred to as hounds right when you said hounds it's been you could refer to any dog and then that has over time shifted to hounds being a specific set of dogs namely hunting dogs um and uh yeah so what's I'm very curious to hear what people are thinking about uh what metadata might go into bibliometrics to make it more accessible a better way for us as researchers to get at Grants and funding and uh not have to be so laborious with power uh literature reviews so with that so probably a bunch of questions so Brenda this is pretty amazing this is really um in very short time not not even half an hour you've you've really uh just totally highlighted this field um I had a couple questions one that came to mind um you know when you're talking about named entity uh recognition and and tagging um are there curated lists curated author lists either provided by uh the indexing services or people um you named a couple of the platforms one biblio metrics which is it's actually a researcher based um and I believe is open source I know it is an ARP it is in our package um and a pretty big platform are there these curated authors lists that could be relied upon because that's a huge a huge gotcha right there yeah I was gonna say so and that's the thing that you know my research has always been so broad that um you know when you're interdisciplinary it can be really hard to I mean there's it depends where you're interdisciplinary right what your other domains are and so yeah I'm not really sure I haven't run across curated list of authors but I'm imagining right that with um that it should be easy enough that top you know main journals would have those lists available fairly readily it would also be intellectual property though right right yeah six one yeah half dozen the other um uh the the other question was actually that I had off top of my head was related to bibliometrics if you ex so what you're talking about is a lot of tools a lot of toolage to to yeah to to um to build this out for a in an emerging field say or a field that hasn't had bibliometrics robustly applied to it like some others might have um so there's you've there's many many different moving Parts in terms of toolage have you explored uh available what platforms what tools exist that you could leverage like just to shine a light on bibliometrics is it something that I know that it's used by people who are doing this to kind of explore and explore their particular Fields it's it's kind of cool that way but um is it does it have tillage that you could use uh for your to to capture the and and and build up a a knowledge base yourself if you know what I'm asking yeah I so just to clarify right um so like um so in terms of uh like Google Scholar let me go back to the but no not Google Scholar but uh scopus right so one of the things on scopus's site right um was let's see let me go back so the secondary documents no this part where you are at the researcher discovery um that that's a kind of a different tool in that it does allow you to look directly at collaborations um gives you but again it's statistics on you know who's it's going to give you something very similar to this type of list um right and so yeah so are we building something you know some of the tools that I've found to be useful still relates to um getting something into a network structure uh so in you know if it if it's not supplied directly from your bibliometric results getting it in is getting into a network structure is its own set of right you've got to get it do that space But once you're there there's tools like graphia is one of the uh free programs that's available to kind of explore it from a from a Network perspective but as far as one there's not yet a tool that lets you run through the whole process nothing I've found I would be curious if anybody else has some stuff that they've found I'm curious about other people's questions I have a question for you Brenda okay good talk this is like part of a whole part of chapter two yes um so when you're trying to do a literature review what do you see as the benefit of bibliometric tools as opposed to just a plain old keyword search in Google Scholar or on a library website I mean why go into the statistical relationships of papers and authors if all you want is a list of things to read well um so great question thank you but so let me I'll respond with this one right in that so yes you could scope it down but you're still going to get getting it down to a manageable level or focused enough um is part of the bibliometric process and the benefit of having something doing something more with it with the bibliometric results right the results help you narrow down what you might want to read the network analysis and the NLP would say if if you're you know if you have 300 papers you might want to start with just reading these three really important ones and decide what that you know as an expert what does that give you to consider well then that begs the question of how do you label what is important because really with bibliometrics you're looking at impact of the paper if you just want to know what's out there a simple keyword search will get you that information and in terms of what you think you should read if it's if you're looking for an overview you don't need a bibliometric result to give you an overview paper or a survey paper you do need a bibliometric result to measure the impact of that paper on the science yes and then how do you factor in the temporal thing if we were studying if we were studying uh generative neural networks and and deep learning that sort of thing and we s and we started our reading prior to 2017 we were going to get one kind of view of the world and then a couple papers started happening in 2017 that just changed everything and we saw just a new propagation of that so I'm wondering to what extent is is that how is those sort of how do these existing tools kind of show those temporal shifts Ah that's actually her thesis so don't make her give an answer ahead of time yeah well to be honest yeah I mean so um you know publication dates um absolutely part of you know the bibliometric results right the the metadata is there on when this paper was published um and you know the idea of course you know is here in the the essentially what I have is the um this last column to the right right where you have you have to identify relationships um you've got to get the topics and then you've got to give weight to those topics and you you know the weight you give can be based on right when they were published and like you're talking John with the what happened in 2017 where you know it was just everything all at once um this is where that would get really fuzzy and I'm not sure how yet that you know what which came first the chicken or the egg and which becomes the you know more highly adopted method uh you know before you can actually get to you know the evolutionary relationship that's but again when you when you've got a whole bunch of stuff and this is going to be really popular especially in NLP that puts out you know 25 000 papers a year um you know being able to tell which comes first is going to always remain nebulous unless we want to start putting in hours minutes and seconds I suppose but you could theoretically get a first occurrence out of an analysis of your bibliometric data correct Brenda um you could get close the the issue being that um again it's going to be limited by the dates really that it's that it's even entered right when it's publication date and uh so you know depending on I suppose if you went by Journal you know if you limited internal you but again that's not really getting at the one topic that you might be after so publication dates made me think of something else they made me think of preprint servers like archive and met archive and the there's two interesting aspects of that one is some of these uh some of these seminal papers are up here or have appeared on pre-print um quite a while quite a few months ahead of their final publication date I think there was example Jim Handler did a commentary on a paper that had appeared I think in the December science um and that that paper first appeared not his commentary but the the preprint of that paper uh appeared in like January of February of last year so it was out there influencing um but the the flip side of those pre-prints is some of those pre-print papers are garbage archive is vetted but uh but archive you just put it there so is it possible to use biblio metrics to BS check to some extent well you see what I'm saying you know if I would tell you that I think I'm meds on the call yes and he's been he's been trying also yeah it's gonna say and he has um quite a bit of you know know-how his thesis work was in uh truth detection right fact checking so he could probably speak to that better than I can uh but yes I agree there's I don't know yet exactly how you know uh how we live in a post-truth world so brother um you've kind of uh I guess slanted this talk towards like um what what papers do you need for a literature review um but I'm also wondering um could this be applied to like say you read a you're learning a new domain right and you read a paper and the first time you read it like it doesn't make any sense whatsoever but it's like an important paper that you need to understand could you use something like this to then um go and be like okay well what papers do I have to Now read to be able to understand the topics that's in this seminal paper actually I mean that I think that would be a really good use of this sort of uh toolkit if you will um to again limit you know I've just recently had to learn a whole bunch of stuff about Prebiotic chemistry and um that has not gone well um because picking up the language is incredibly difficult I am not a chemist at any stretch and so yes so having something like this that yeah points you directly to if you can pull the keywords or you know authors from the seminal work that you start with um yes you should be able to limit what or direct yourself towards the next best paper to go to so so I'm just wondering um so that something like this means this doesn't obviously exist yet and um so so what is your goal to actually create this or a piece of it or um so um have you actually done any of the NLP parts or um like how far have you gone in the workflow are you trying to get through the whole workflow where the end result would be this literature recommendation system or are you planning on doing kind of like just part of the semantic analysis piece of it what's your goal for this that's a great question so ultimately right so I have used a test case it was a high energy physics theory uh data set and have basically ran through this whole process and ultimately what I'm leaning towards is being able to um look at the evolution of a scientific field so um so yeah basically that's it I want to wind up with some way that you could easily go and say okay so how did um how did these theories evolve um you know so we should be able to trace back where that theory started in literature and then ultimately see how it uh essentially densified if you will right as it became more popular it would it would you know you would see an upward Trend um and whether or not that falls off or uh you know exactly how that evolves and what the effect is who's tied to that is really my interest and that's your deliverable to create something that can do that and describe a it's it's basic basically what I'm tasked with doing is creating a descriptive analysis of uh of a research domain essentially a scientific domain interesting good luck yeah thanks gonna need it I have a question yes uh so first of all great talk thank you I can't wait to try out some of the tools you mentioned myself um I'm curious though because I'm kind of new to this area so in bibliometrics when you're looking at a citation for the paper are all citations created equal I know you talked a little bit about how many times they are cited they use themselves but do you look at because I feel like there's a difference between the citations you put in a paper like there's some citations where you're pulling on some definitions from earlier work others where you're looking at related work that you're trying to compare yourself to and maybe other papers that are a little bit more core to like the main thread of the evolving research are those sort of features uh explored biblio metrics and that's you know so kind of at the core of bibliometrics is the whole citation analysis and um there is a ton of literature on just how biased this uh you know that whole genre itself is and so there's a lot of problems with citation analysis just and the way we cite papers um is a is problematic um so so there's no so so it's a faulty system to begin with is what I'm trying to say and there's no real further metadata necessarily tied to it so figuring out why someone uh cited a particular paper may not be clear and I'm not sure how I would get at whether or not okay yeah thank you yes any other questions I love the idea that not all citations are created equal [Laughter] uh so early on I think you mentioned Bradford's law or yes someone's law right I didn't really understand it so so you have a paper and you read that paper and that paper has a bunch of citations and is Bradford's law or I guess in your example you read like five papers and each of those five papers have a bunch of citations and so is Bradford's law is don't even bother reading the citations from those papers because you're not it's not it's not worth going down that rabbit hole so yeah so though so it's the it's basically a law of diminishing returns if you will right so basically um what it what it is is it's much more focused on the at the journal level so it says you know basically if you know you know in um I couldn't even give you five top journals in data science but you know say you had five top journals and you know and out of those five journals you get in in a month there's actually 12 you know say there's 10 articles that are actually worth your time to read um in order to find 10 more you have to search 10 more journals it's a doubling situation right so you would have to search 10 more journals to find five more articles or ten more articles and then again then to find 10 more then you've got to do another 20. and it keeps increasing exponentially so there's it's like I said it's just basically a law of diminishing returns so it's so the the issue with the journals is that so when you go to bibliometrics they're only going to source those top five journals because they know that researchers won't go beyond those top five because it doesn't make sense time wise for you to continue searching so it's kind of an area area under the long tail curve is what you're sort of saying you have to just keep going out there um the the the middle law so so that was the the it's Bradford's law right Bradford Bradford's law yeah and then with lochte's law was the uh lot cuts and these are both derivations if you will of Zips law zipf um and um so latkes is is right it it's a frequency right frequency of publication and so with authors so basically the what it describes is as you know as the there's an increase in the number of Articles being published the author's publishing those articles actually decreases so you have an increase that's a terrible way to describe that um so basically there it's a it's a given ratio right of the number of uh authors publishing to a sing a single article right so the ratio doesn't change a single article or a single topic yeah I think you meant topic topic yes single topic which kind of makes sense right because like when when a author first publishes on a new topic they're going to be the only author publishing on that topic and then um so their ratio of publishing on that topic increases but then as that or or as high to begin with um but then as the topic becomes more well known more people are publishing on that topic so then um the authors percentage of articles on that topic uh starts to decrease as more people start writing about that topic unless their name is Robert Hazen in which case they're on every part what's that I say John and I have type Robert Hayes and a few too many times in a d space well and Ahmed Alish and Anna rude and SM Morrison um but okay so but that kind of so laka's law it kind of characterizes the whole hog piling on a topic sort of yes thing yeah so so it's again so again skewing the data per se if you will you have to be aware of it so that you can control more right yeah Bradford's is a little less intuitive well it in a way it is in a way it it isn't I mean I think that so you had asked a question earlier on kind of how to use these tools and that sort of thing and what I was reflecting on you know the the olden days you would try to find um as you commence a literature review on a new domain you tried to find those golden papers a few papers that were very citation rich uh whether they were review papers or just seminal papers but they were citation rich and then you kind of took you use those as almost like a textbook uh and and a reading guide to exploring that cracking open that field and and that would create kind of uh a network you would construct this network of other of other literature following following those the the net that was introduced by those few handful seminal papers um now we have these tools to to help you do that rather than microfiche and other cranky hard catalogs Librarians that requires talking to a person but yeah and we happen to have a new wonderful Eye Brand if you haven't met him he's worth it so you so you asked a really good question when you've when you finished your the the first part Brenda is what metadata um really I I guess what would be on your wish list for the you know what metadata do you look forward to creating that helps really um enable this work to be done in a new field well I mean so there's generally um somehow I would like to see in the metadata itself um more information about the paper so generally right there they allow for authors to supply keywords um when that paper is submitted what's that smells same peasant stem chart a little whip up when we first met ah okay sorry I've had something in my ear should have been that's really weird yeah so yeah I I don't I can't there's too many possibilities okay so if I had to limit it to you know two um top metadata Fields I would love to see the uh paper information and that for me would be you know yeah my brain's going from really big to really small here's a way that a different way you could answer that you know Brendan some of the conversations you and I have had you've talked about the need to sort of progressively NLP into the paper to get more depth talk maybe talk a little bit the difference between uh you know title you know getting in from metadata from the extracting data from the title versus the abstract whether or not to actually consider keywords introduction you know those can once you maybe talk through um progressively how you can get more value well so you know papers have right we all know this from reading them right papers have standard sections right so we have the introduction the methods section um the results the discussion um and you know ultimately the references right now you know we keep good track of the the references that we could actually have access to you know individual sections of the paper um you know in the introduction generally we cover all of those you know those historical things if you will and areas that you know that we're bringing together to to do our research so there's going to be a lot of citations potential in in that particular area of your paper generally um and so with that you know then if you can actually you know if I'm doing a literature review I just might need just the methods and if I could get it just the methods boy that would help me just streamline you know exactly what I what are the popular methods what who is doing you know I could actually say who's doing which particular method what maybe I could even pull out you know what are the particular um equipment that they're using right uh to do these you to apply these particular methods um you know so there's all of you know just breaking down the paper itself seems to have a whole bunch of potential for you know mining you know that kind of of you know just comparing an intro to an intro to you know more likely comparing a method section to a method section um would really be just eye-opening for what you could find um I would think the different kinds of papers I mean uh in AI paper and a um bioinformatics paper paper and you know on you know from the tile Consortium on on Alzheimer's and stuff there's those two different worlds but uh but yeah the the latter is very highly constructed they're very dictated in what the structure of that of their what they write is so right yeah one of the problems of of NLP you know just even as an entire field is that that you know there's not a one-to-one correspondence with you know word to meaning so you know until we have a one-to-one correspondence how accurate you know NLP can get is always going to be somewhat nebulous they're not really going to be able to be real accurate and I was going to say you just um scanning for and comparing methods or scanning and comparing uh literature searches and papers is not going to be useful because you pull out you you destroy the context in which all of that stuff was written right so at that point you're really just looking at a word search and you still need to apply a lot of human brain power to it so if you were going to do something like comparing methods from across different papers you would that's that's a whole different dissertation project than you would definitely need to combine a couple of MLP methods there well I and I would say too also to that is that you know but if what I'm doing could actually lead to this space right and that if you can actually drill down to finding you know the papers that are very relevant then being able to just pull just those methods and find the common methods that are there can be can help provide you know fodder for you to you know discuss which method you would use and why so this works within a discipline what if you're trying to do something cross-disciplinary this is where you know you see semantic shifts here but on the slide but one of the things that's that's really becoming you know an issue as we know is you know science itself is becoming more and more interdisciplinary cross-disciplinary multi-disciplinary if you will and with that you know what's happening is um you know one particular domain defines a particular process or a particular uh you know item in it with a set of words that you know the the neighboring domain may not use any of those words at all to describe it um and so you know and not to mention you know that you have those situations where when a tool was originally invented um it was called wasn't called you know a a mass spectrometer when it very first came out you know it's that sort of thing so you know tracing that back along with these other um elements of that and connecting those across disciplines is is a definitely an N hard problem so yeah not not clear on that on how to get at that yet although Dr groupie who recently graduated did some great work on semantic shift and yeah so hopefully there'll be more work in that area coming soon right do we have any other questions out there with that let's see guys I think we're supposed to use our little hand wavy things for a WebEx we have a lot of practice there we go all right and I'm going to stop the recording

Original Description

Bibliometric methods, such as citation analysis, have been the primary means of studying the development of scientific disciplines. However, these methods have limitations, such as the inability to capture relationships between research concepts or identify emerging areas of research. This will be a discussion about bibliometrics, overcoming its limitations, and what possibilities exist if we explore the various relationships within scholarly literature using natural language processing and network science techniques.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Playlist UU4rjm_R9sgRNvv9QsgH8LDw · Tetherless World · 23 of 40

1 TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
Tetherless World
2 TWed Talk: Neha Keshan on "Stress and Machine Learning"
TWed Talk: Neha Keshan on "Stress and Machine Learning"
Tetherless World
3 TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
Tetherless World
4 TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
Tetherless World
5 Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Tetherless World
6 Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Tetherless World
7 TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
Tetherless World
8 TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
Tetherless World
9 IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
Tetherless World
10 TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
Tetherless World
11 TWed Lightning Talks Fall 2019 (11 Dec 2019)
TWed Lightning Talks Fall 2019 (11 Dec 2019)
Tetherless World
12 TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
Tetherless World
13 TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
Tetherless World
14 TWed Lightning Talks Spring 2020 (29 Apr 2020)
TWed Lightning Talks Spring 2020 (29 Apr 2020)
Tetherless World
15 TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
Tetherless World
16 TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
Tetherless World
17 TWed Lightning Talks (Fall 2020)
TWed Lightning Talks (Fall 2020)
Tetherless World
18 TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
Tetherless World
19 TWed Lightnining Talks: Spring 2021
TWed Lightnining Talks: Spring 2021
Tetherless World
20 TWed Lightning Talks (Fall 2021)
TWed Lightning Talks (Fall 2021)
Tetherless World
21 TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
Tetherless World
22 TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
Tetherless World
TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
Tetherless World
24 TWed Lighting Talks Spring 2023
TWed Lighting Talks Spring 2023
Tetherless World
25 TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
Tetherless World
26 FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
Tetherless World
27 FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
Tetherless World
28 TWed Lightning Talks Spring 2024 (14 Feb 2024)
TWed Lightning Talks Spring 2024 (14 Feb 2024)
Tetherless World
29 FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
Tetherless World
30 Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Tetherless World
31 Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Tetherless World
32 TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
Tetherless World
33 TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
Tetherless World
34 TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
Tetherless World
35 TWed Talk: Thilanka Munasinghe (26 Mar 2025)
TWed Talk: Thilanka Munasinghe (26 Mar 2025)
Tetherless World
36 TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
Tetherless World
37 "Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
"Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
Tetherless World
38 TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
Tetherless World
39 TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
Tetherless World
40 TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
Tetherless World

Related Reads

Up next
Welcome to the Next Temperamental Era
Charles Schwab
Watch →