Lecture 18: Case Hx: Cancer Diagnostics

MIT OpenCourseWare · Beginner ·🔍 RAG & Vector Search ·3y ago

Key Takeaways

This video lecture covers the application of genomic medicine in cancer diagnostics, focusing on the use of microarrays, gene expression analysis, and retrieval augmented generation (RAG) search to understand the molecular determinants of cancer outcome and develop targeted therapies.

Full Transcript

so um i i come with this actually originally from a pediatric oncology perspective so i'm going to start by giving um examples of a couple of patients that i saw um in the jimmy fund clinic at the dana farber it was sort of typical so the first patient was a nine-year-old girl who presented to a pediatrician with uh she turned these up with a fever and bruises um she got a blood test and a bone marrow test that revealed that her bone marrow had been replaced by acute leukemia cells um acute lymphoblastic leukemia cells or alol and so she was enrolled on what was a standard chemotherapy protocol which was like nine different drugs in rotation and combination she entered remission in three weeks and is still alive and well and then a few months later there is a second patient this kid happened to be a boy about the same age same presentation same diagnosis acute lymphoblastic leukemia was enrolled on the same treatment protocol so got the same drugs same hospital setting so it was as close to a controlled experiment as one could do in a human being because of course response to any kind of therapeutic intervention result it's not only a measure of the treatment itself but how it's delivered so that was controlled but this patient didn't respond unfortunately and died about six months later so the overarching question that i'm sure you've been addressing in this course also is is generally how to understand this clinical variability and to discern whether there's a molecular underpinning to that variability um so we're particularly interested in this patient who well responded particularly well and so did what pretty straightforward which is to do a standard chromosome analysis carry typing experiment using some molecular techniques uh known as fluorescence and situ hybridization details don't really matter except to say that while it wasn't apparent at the routine morphological analysis of the chromosomes if you look molecularly it was clear there is a translocation between chromosomes 12 and 21 that fused two genes actually two transcription factors one called tel one called aml to make this tell a ml1 translocation and it turned out that while it hadn't been previously discovered actually the the majority of known genetic abnormalities in childhood leukemia the the lion's share of those actually have this translocation and of those patients in retrospective and now prospective analyses about 90 of those patients survive uh turned out the patient number two uh had a different translocation a 922 translocation that fuses two genes bcr unable that you've probably talked about and it's known that these patients given the same therapeutic regimen only have about 10 percent uh survival and so these these are molecular tests done at the time of diagnosis so what i think this really sets up is which is now perhaps generally accepted but was really emerging at this time was the notion that cancer is a genetic disease it's one second and that the outcome clinical outcome is predictable based on molecular determinants at the time of diagnosis as long as you know what to look for yeah is before greenback turns out that even with gleevec which targets bcr able in these patients it doesn't it's not particularly effective uncharacterized multiple balanced translocations that identify a single oncogene at a breakpoint so we started thinking about alternative ways to try to think about molecular classification of cancers more generally and this is just the obvious kind of experiment two biological states that could be clinical states or biological states um that you collect some kind of genomic information on like microarray data and then have this pesky problem of actually how trying to figure out how to interpret the patterns that emerge and i'll say quite a bit about that part because it's tough um have you talked about the non-analytical but sort of just sort of laboratory aspects of microarrays at all let me just spend one slide talking about this so all of the microarrays are based on the basic principle of somehow labeling mrna or its derivatives from a cell and hybridizing to probes on some sort of a solid support whether it's a microscope slide a silica wafer something um they're cdna rays and they're oligonucleotide arrays and you can either make them yourself or you can buy them commercially cdnas generally take the form of two color hybridizations where you simultaneously hybridize a test sample and a reference or a control whereas the oligonucleotide arrays generally use a single color this is completely and entirely a historical artifact of how these things were developed there's nothing intrinsic about cvna arrays that um [Music] makes them require two color hybridization the only reason for using two color hybridization is if the quality of the raise is so bad that you need an internal control for every spot otherwise you can't interpret the data and such was the case for the earliest cdna arrays which are the first ones to be made and i would venture to say that cdnas in general are going to become obsolete very quickly in favor of oligonucleotide arrays probably commercial ones the genome is finite once you have a genome represented and identified probably not much advantage to making your own and most of the arguments that say oh we can make them much cheaper making them ourselves are usually uh sort of enron type accounting uh uh practices where you don't sort of well we didn't count the cost of all the people involved in making these things we didn't include the fact that we spent three years trying to figure out how to do these and they didn't exactly work when you really get right down to it usually buying them is cheaper and they require quality i should also mention that there are non-microarray based expression profiling methods like serial analysis of gene expression or something called mpss that are transcript counting based methods where you use essentially dna sequencing to enumerate the precise number of copies of a given transcript in a given cell and the proponents of this method um say and they're correct this is the only way to really know how many copies of a particular transcript there are in a particular cell i would say that's absolutely true but it's actually not that interesting because for most biological questions it actually doesn't matter how many absolute copies of a transcript there are it's not that useful information what's useful is to have some kind of comparative experiment that tells you there are more with statistical significance more in this sample than that sample and rarely do you want to know that there are 682 copies you know but it is true uh but i would say that the the the low throughput and cost of these counting based methods greatly outweighs the benefit that you get from this absolute counting based method [Music] uh well copies of the transcript not of the gene so you could say we sequenced we identified a million total transcripts in a cell of which 684 came were transcripts from encoded by gene x right so that that can be useful kind of quantitative information but not so useful to i mean these these experiments are still you know many thousands of dollars per sample um [Music] when it comes down to like a pharmaceutical company trying to figure out the mechanism of the disease and actually how to go in and develop target drug molecules that might interfere like i'm assuming that this technique would be useful for example if you do if you do one of the above to oligarchy and you don't know the quantitative number of transcription present you just know sort of qualitatively more or less compared to disease state or not yeah um if you know if you know a number and does not give you the ability to make speculative speculation about you know how those transitions are being processed and whether that's an important process in the department i would say it doesn't actually because if you tell me you know if god tells me that there are 684 copies of a given transcript i don't know what that means i do know what it means that if there are you know five times as much of a given transcript in a disease state compared to normal state i can at least say that having one-fifth the number of transcripts isn't sufficient to give you the disease state right whereas if you just give me an absolute number without a comparative sort of thing um i don't know what that means and i'm not sure that saying there's 600 in this state and three thousand in that state to me that's not really more useful information than knowing relative abundance how what the resolution is in terms of common differentiation that's like you get it usually you get an exact number for the bottom reference but for the top you get a sufficiently resolved difference between samples that you can say that's right that's right you know 104 that's right and i would say for most biological and clinical applications knowing that the ratio is one to five gives you about 95 percent of the information that's useful compared to saying it's 500 versus 100. there might you might be able to think of some special experiments where you really want to know the number but usually it doesn't help you very much in my opinion there are all sorts of sources of variability that we won't discuss except there's only one that matters and that's the um biological and clinical variability that goes into these experiments with a lot of people that spend a lot of time handling over these various technical things and none of them make any difference really as far as i can tell because they're overwhelmed by biological variability so making microarrays that are somewhat more accurate or precise in their measurement won't actually make that much of a dent in the problem because that variability is exceeded so tremendously by biological variability that's all i'll say about that so let me give some examples of applying this and i know you've discussed these general methods before but that's okay so here's an experiment where we're interested in the differential clinical outcome of children with a brain tumor called medulloblastoma you didn't just have you discussed this particular example okay so we had 60 pre-treatment biopsies of patients with this brain tumor tumors were biopsied and we knew the long-term clinical outcome of these patients whether they uh survived from their disease or they died despite therapy they had the tumors removed and were treated with chemotherapy and radiation and we knew that some of the patients survived and some of them did not and based on that we said well perhaps there are two classes subclasses of disease so we clustered the data this was across i think 6800 genes on a on a microarray so each dot represents a different patient sample say well if they're two classes let's cluster into two groups um because we have this clinical suspicion there are two classes and of course if you ask the algorithm to cluster into two groups in this case it's something called a self-organizing map but it doesn't matter you could take the two major branches on a dendrogram from a hierarchical clustering dendrogram whatever you get the same picture you get two classes and the patients are shown there about equally distributed and now if we fill in the labels of the patients that is whether they're survivors turned out to be survivors or not survivors you get this picture where um you don't need a statistician to tell you that there's no correlation between this class structure and survival so what do you take what do you what do you think of this experiment what can you conclude what do you do with this how are you going to rescue this are you saying that there's no signature survival right so i think you're both hitting on the the relevant point here which is that the unsupervised learning clustering algorithm found some structure dominant structure that made these two classes they just don't happen to have anything to do with the question that we're interested in which was one of survival and so that gets back to this basic notion of two general approaches to to data analysis which you've probably discussed but i think often gets confused so unsupervised learning which is not exactly synonymous with clustering but it's a reasonable first approximation or supervised learning uh classification so here you're interested in finding dominant structure defined only by the intrinsic gene expression patterns in a given data set irrespective of anything you happen to know about the samples such as their clinical outcome here you're saying whatever maybe there's some other interesting biological structure but i'm not interested in that right now i want to know whether it's a gene expression pattern that's correlated with the thing i care about in this particular example [Music] outcome so we take the tank the same data set the same matrix of data samples by gene expression values and now apply um supervised learning approaches this happens to be a k nearest neighbor classifier again it doesn't make any difference what you use here and this happens to be an eight gene model you make a you classify the samples using a level and out cross-validation approach so that you don't you attempt not to over fit the data the model to the initial training set and then you ask well of the two classes that are that are predicted how do those patients actually fare and here is a survival plot in terms of months months of survival for those patients who are predicted to be alive versus those that are predicting to be dead so what do you think about this is this significant how would you so if you look if you look in a you know basic biostatistics textbook about how to calculate statistical significance of a kaplan-meyer survival curve such as this is they'll tell you to do the log rank test and if you did that you would get a p-value that if you've looked at a lot of these things would sort of match your intuition for this degree of separation okay um so it looks something like that is that reasonable right well i mean that looks quite significant but you should ask well how come this model has eight genes for example how do you choose that number well it's quite easy we chose that because it worked the best worked better than six worked better than 10 or 50. um so we have to pay some penalty right for for overfitting a model potentially to this particular data set so the ways that you could then really test the significant the statistic statistical significance of this model would be to apply it to another data set i that's the gold standard but short of that a reasonable thing to do to better approximate the significance is to take into account the fact that there are a number of parameters of this model that were optimized to fit this data set and you can do this by doing a permutation test where you don't scramble the gene expression values themselves but you randomize the class labels in terms of are the patients alive or dead and you go through the same procedure of attempting to build an optimal classifier and choosing including choosing the number of optimal number of genes to ask if you really try hard with these machine learning methods how often can you make a classifier that works as well as this one does and when we did that a thousand times nine times out of a thousand we could do this well or better so we estimated the significance of this model here which is still you know decent but you can see we took a couple of a hit of a couple orders of magnitude on this p-value so if you had had a nominal p-value of 0.05 or something right that result would entirely vanish when you appropriately attempt to correct for such multiple hypothesis testing and this is independent of what particular classifier you use so i would say much of the literature and you know everyone's figuring out how to do this sort of as we go along but much of the literature and worries about failure to reproduce an initial model are due to the problem of over estimating the significance of an initial model because of these overfitting types of problems so let me just make a couple of general comments about supervised learning and some of them may seem obvious but turn out to be actually problematic and this is one example so the first step establish the class labels of what you're trying to classify um so in one of our first experiments we're trying to classify the two different basic types of acute leukemia acute lymphoblastic leukemia or acute myeloid leukemia the way you build a classifier is to choose examples of one class examples of another class and then find gene expression patterns that are correlated with that well who's to say that we have these right right what should you use as the gold standard for these things it's not always obvious particularly for the very things that we want to build better molecular classifiers for because the current clinical diagnostics are so poor it doesn't really feel like a good gold standard to go back rely on the clinical labels as the gold standard uh and so this is in general a real problem for survival studies if you force the question into a simple two-class problem survivor or non-survivor well everyone's a non-survivor at some point so at what point do you declare that a patient is a survivor of their tumor for example that requires some judgment call in terms of which bin to put the samples into and i would say that much of at least our effort and time has gone into trying to figure out how to get this right and there are some approaches that one could take to not have to be so rigid on how you how you assign these labels but it's a challenge so the second general um step in in making a classifier is selecting the features that you're going to feed into a model features in this case being genes the details don't matter their whole long list of ways that you can rank genes this is a simple one that is based on the mean expression level in the two classes and their standard deviation it's a relatively unsophisticated way to select genes because it assumes that there's a uniform behavior of the these marker genes in the two classes which in many cases is not at all uh the case there are many other methods as well and then i'm sure you've talked about these at length so i think maybe i'll skip this that you then take these things and classify so for unsupervised learning there are all these methods and i think they basically don't matter that either for for clustering or supervised uh machine learning methods if you get a result that that is obtainable with only one magical you know one person's really special algorithm i would worry deeply that there's a problem uh with you know there's there's an information leak somehow or something's not right uh because at least in my experience when there's really biologically or clinically meaningful structure to be found you can find it with a number of different approaches in fact that's a reasonable sanity check to make sure that you can recover structure whether you're using machine learning or unsupervised uh clustering algorithms that you can recover it with multiple different methods there are some examples to that that can be interesting but on the whole um i'm confident thing it doesn't really matter and it's really the input to these data sets that matters the most that is are you really sampling the diversity of for example the disease process that you're studying with the samples that are in your initial data set what's more challenging is how do you evaluate the output of these clusters in terms of their biology in terms of really knowing how robust the structure is given any algorithm uh and then how do you actually know when you having seen a given structure once how do you actually apply it to another data set to know whether it's you see it there as well it's not obvious so let me give a couple more any questions about that general stuff i wasn't going to say anymore because i know you've covered it let me give a couple other examples of applying these principles to some some data sets did you talk about this one no okay so here again focusing on childhood leukemia most kids with childhood all respond to chemotherapy i told you about this subgroup of bcr able patients it does not another group that does not respond well are infants less than a year of age who generally don't respond it turns out that most of those patients have translocations into a gene called mll um and so but it's clinically of interest because these patients don't respond to conventional chemotherapy and this just shows you that using standard clinical criteria they're hard to distinguish so what if we take conventional all samples these infant mll rearranged leukemias and some aml the myeloid leukemias and we apply an unsupervised learning approach uh this happens to be principal component analysis but could be your favorite clustering algorithm what do you see here three different classes why do you say that jose speak up you might not know okay so right here so yes you see three classes but only if you have the right the colors filled in right so if you imagine this is just a group of of leukemias you might get the sense that there you know was something going on over here but if you imagine these all black it's not so obvious maybe no this is completely unsupervised so that's that's the first point is that these things often look clearer when you actually impose knowledge on them even though the structure here is done an unsupervised way you get this sort of impression that's a really clean result if you superimpose knowledge afterwards okay that's the first thing but let's say yes there are three classes and i think you can appreciate that one question was when maybe these infants with the ml rearrange genomes in green maybe they don't respond to therapy because they're babies and you know that this is a metabolic post-metabolism problem and their leukemias are the same as the conventional alls shown in in dark blue this would argue that it's actually not the case that they're fundamentally different leukemia is this helpful [Music] so it's sort of helpful maybe from a taxonomy perspective but does it tell you what to do for these patients right so what if you wanted to gain some biological insight into what was different about these ml green infant leukemias what might you do you've got these data you see that they that those patients define a different class what could you do anybody have any idea what would you do with this the genes whose weights explain the most this separation that's right so you could that uh you could do that it turns out in this case there are a lot of genes that actually have relatively equal weight so you still have a large list and the three principal components don't perfectly separate the classes so you could go back and say well now i believe that these mll leukemias are a distinct entity so now you could go back that would be reasonable no that's a reasonable thing to do but the other way you could do is to say well now this tells me that i believe that these ml leukemia is our distinct entity now let's use supervised learning types of methods to identify the genes that are most correlated with the class of interest for example high in the ml class versus the others that would be a straightforward thing to do right so you could rank the genes according to that distinction so we did that and did what i think yeah the first you case be comparing the between the different classifications uh the difference that the different genes can expression over there and the second phase which is no well no i think it's more that if if you didn't have these colors to look at and you said ah there's some structure here i don't know what it is what's the biological basis of this structure looking at the weights of the genes that are driving this distinction would be a reasonable thing to do in this case we had a specific question are these leukemias unique or they add mixed with the others having determined that they are unique it's a little bit cleaner to say all right let's use supervised methods to find the genes that distinguish one class from the other of course if you had perfect separation it would be it would reduce to the same experiment but because it's imperfect um there's some advantages to using class labels here i should mention also that you see this blue guy here sitting in a sea of green so this is a patient that based on gene expression one would predict to be ml rearranged but the clinical record for this patient said he was not but when we went back and actually looked at this turns out that there was a missed translocation into the mll gene that you could recover by fish so you know there's this is not a public health menace you know uh diagnosing these leukemias properly but there are examples of of missed diagnoses that i think can be i think looking at these multi-parameter gene expression readouts can serve as a a unifier an integrator of lots of upstream genetic activity and so i think the power to detect those upstream events is going to be higher when you look at some downstream pattern such as an rna pattern as opposed to developing specific tests for each of the individual genetic abnormalities that could cause the same phenotype because in the end all you care about is is knowing whether the molecular program has been activated but so if you rank the genes according to this distinction and just start at the top of the list uh here is the gene that was top of the list of twelve thousand six hundred whatever that were on on the list um and anytime a tyrosine kinase rears its head in a cancer classification cancer biology experiment you pay attention to it particularly given the gleevec story right um so what do you think about this i tell you oh look at this expression the rna level of a kinase so receptor tyrosine kinase called flip 3 is um characteristically high in the mls compared to the others what do you think of that in terms of therapeutic potential therapeutic significance so yeah what would you do with that or maybe this already popped up because these kind of slower right so we define this list by virtue of the fact that it's high in the mls compared to the other two combined um to a flip 3 inhibitor so you want to treat patients with a flip 3 inhibitor well that's not an fda approved drug so you can't do that [Music] so there is so there there are so the quite the underlying question is the hypothesis would be that ml leukemia cells are dependent on flit 3 kinase activity for survival if that's not the case then you don't care then unless that's the case the over-expression of this thing is totally irrelevant from a therapeutic perspective right so you could do that genetically for example using rna interference to knock down the expression or you could do it pharmacologically if there was a compounding development not yet a drug but a compound in development that inhibits kinase activity and so that's what this experiment is [Music] you can do it in you can't do it in in vivo in a person but you can do it in human derived cell lines so clinically and that would not be like you know hypothesis that you want to go down that path that's right now you can you can make the argument that well you know doing these things in cell lines and mice that's not real disease and so i don't care what any of this stuff shows but still if your hypothesis is that a given gene the overexpression of a given gene is important and you do the experiments to ablate the expression of that gene and nothing happens that should deflate your enthusiasm a little bit okay so here's the experiment though here now taking patient derived human cells that have been engineered to express firefly luciferase gene so that they glow and you can monitor in vivo tumor burden so here mice that on week one you inject in the tail vein infant leukemia derived tumor cells and you see over four weeks time the amount of luciferase activity increases as the cells grow and the mice start to die around week four and here is uh a cohort of mice also injected but treated with a drug once a day by mouth that functions as a flit 3 kinase inhibitor and you can see that the development of the leukemia is significantly obligated which at least to our first approximation validates the hypothesis that flip 3 overexpression is not just a diagnostic marker of this class but it's actually a potential therapeutic target and so based on this and some other preclinical data um the clinical trial that you wanted to do with a flip 3 inhibitor is is being planned to treat patients so the the cells are infected with a retrovirus that contains a cdna for firefly luciferase gene so that if you inject these mice with the compound luciferin they will emit the same enzyme that fireflies do and they will glow so usually this is done in vitro in test tubes but here you introduce it into the cells in the animal so that you can monitor what you used to have to do would be to inject a whole bunch of mice kill some of them here kill some of them here and then like examine the bone marrows to evaluate the progression of the disease what's nice here is that you can follow a cohort of mice non-invasively yes dumb questions when you actually look at these mice can you tell that they're fluorescing no no uh no you need to use a special device that can measure i think in the near infrared um range there are green fluorescent protein mice that actually do glow and you can tell that they're green the mice that you use are they um like immune like rag knockout if you don't have a massive immune function yeah so you have to do this in immunodeficient mice so they don't uh reject the human tumors and how do you account like does that factor in your determination of like the degree of you know the potency of the spine of the tumor cells and whatnot because of just sort of the fact that the images you can direct attacks against and so when you're when you're considering these experiments you're saying okay i see this you know spread across the entire mouth in this level how do you factor that you don't you factor that in by saying that there are many things that are occurring in these models that don't recapitulate what happens in mouths most people don't get cancer by having intravenous injection of a tumor into them you know most patients have an immune system you know so i think it's just one of the limitations that yeah it would not be worth the time and expense to have a drug development project around every you know little inkling that comes out of a microwave experience and you do something even though it's deficient in many ways and you've hit on some of them to say is this interesting or is it not i think not yet at the point where one can do this entirely computationally and have any kind of you know confidence that being said these so-called xenograft models where you put human tumor into a mouse model are not particularly predictive of efficacy of a drug in a human clinical trial but in the absence of anything better it's still what most people do first is more robust and like if there's no effect on the xenographs then you're really a loser if you go to the human no it's actually more if anything at the opposite that if you show some activity in the xenographs you often see activity but failure to see activity failures the activity of xenographs is not particularly particular for molecularly targeted therapies where you it may be that you can show in the mouse that you've really shut down the pathway let's say you've inhibited flip 3 completely so some people are start drug companies starting to use xenographs in that way to say all the mouse is is a test tube so that i can ask have i inhibited flip 3 enzymatic activity yes or no if i have and i believe that flip 3 is a good target i don't care whether the tumor's actually shrunk or not i'm going to bring it forward to to clinical trial but you need something to convince you that the target is reasonable yeah yeah i think you see them there because the the cells they're injected intravenously in the tail vein but they home to the bone marrow and you're seeing large bone marrow cavities which is why you see them kind of over the flank there i think i'm going to skip this okay you talk about this no good um so when you do these experiments the data usually present you with two you either have one of two problems after you do all this appropriate correction from multiple hypothesis testing that i told you about either despite having done that still have a list of genes that's too impossibly long to bring biological understanding to or you've corrected away everything and you have the impression that there's actually nothing that is differentially expressed in your two cases and so let me show some of the more recent approaches to dealing with this because it's really substantially changing our thinking about how to how to do these kinds of experiments so this is not a cancer example but a diabetes experiment where there were patients who were either adult patients who either had type 2 diabetes or they were normal as defined by having a normal glucose tolerance test and they underwent voluntary skeletal muscle biopsies under eu glycemic clamp 18 of these patients 17 of these patients it's a simple two class problem right get the expression data to identify those genes that are differentially expressed in these two classes do the appropriate permutation testing to make sure that you correct for multiple hypothesis testing and here's the result nothing meet significant so even so out of the 20 000 genes on the array even the top-ranked gene doesn't meet statistical significance possible that this is the case but the question is are there other ways that you might go about recovering a biological story here and so the way that ramsay mutha and graduate students arvind subumanian took to this was to define groups of genes or gene sets whose activity as a collection of genes could be interrogated in these data sets and we could have a rich discussion about how does one define such gene sets you could do it based on the literature so ask zach what genes are important in you know some pathway that he knows something about that could be a list or you could say uh we don't trust zach about you know that sort of that'll bring zach's bias we're not interested in that or anyone else's intuition let's just experimentally derive lists of genes by one way or the other perturb cells get the gene expression change and that makes a gene set okay and you can collect as many of these as you could stand for these experiments we made 150 of these gene sets okay before you go online twice on this slide oh yeah yeah he's enriched here he should be uh so then how do you do this so first thing you do is rank all the genes on the array one through twenty thousand according to how well they're correlated with the distinction i already told you this top one even that one doesn't meet significance in a single gene and then you interrogate each of these gene sets and ask are they enriched and so here would be an example of a hypothetical gene set each gene in the gene set of a dozen genes or whatever that is not enriched towards the top of this rank ordered set list whereas here is a hypothetical gene set it's not perfect right but it's non-randomly distributed on this rank ordered list it's enriched towards the top you see this as a similar operation to um the following there's a bunch of programs out there look at forgiving microarray result at gene ontology and say what classes of of gene and today are over-represented and they give them in the set uh yes that is so you can define these gene sets based on a gene ontology annotation that's an example the important part is to make sure that you appropriately correct for testing all these gene sets so now instead of 22 000 genes we have 150 gene sets but that you should think of like 150 hypotheses right so you should do the same sort of permutation type of testing and say if i randomize in this case the diabetes or normal distinction is my favorite gene ontology class still enriched and that's what some of the current approaches that kind of annotation don't don't do and so you can codify this in something called the kolmogorov smeared off statistic it doesn't matter you can come up with an enrichment score for these things if you do that in this example you essentially get one gene set which meets no just which gets quite high statistical significance for a set of genes so how do you reconcile this how do you reconcile this thing and this thing how could that be that's my question so the first thing you're saying that you didn't pick up any difference in the expression of a gene by gene basis basis but then when you group a couple of them together all sudden there is a difference right so how could that be how could that be the microarrays themselves the precision of these arrays you know is not so fantastic and so you can imagine that there's a subtle signal on a gene by gene basis it's difficult to detect it but if you consider the coordinate regulation of a group of genes all in the same direction as a group this might be quite striking and this is shown right here which is really quite amazing when you think about it so here we'll look at the mean expression level of all the diabetic patients versus all the normal patients all the genes on the array are shown in gray and so as you would expect there are no outliers right there's nothing really way off a diagonal if they were those would show up as as single genes that were differentially expressed of course you could have one massive outlier that could screw you up with looking at the means but still you get my point and here are these oxidative phosphorylation the gene set that was defined by those genes that are involved in oxidative phosphorylation and you can see that with only a few exceptions they're all lined up just below the diagonal their change in gene expression is only about 20 percent compared to normal but it's all in the same direction so 20 percent change in this number of genes is is quite quite significant so this are you getting that let me try to because it's an incredibly important point um the chance if you look at any given dot of being what it means to be one on just near the diagonal one side of the other you might make a much story or about the fact that it's on one side of the other and then just by dump luck they're all on one side of the diagonal if they hook that up that's going to be incredibly unlikely and so each individual gene being on one side of the diagonal the fact that all genes that you have assigned beforehand would be other types in this case oxygen phosphorylation they will end up on one set of diagonals that's hugely unlikely the fact that you can just like dumb luck put them all on one side of the diagonal so it's less tied to the enrichment levels that you find that the given that you've said these are genes that are should be related to their interaction yes that's right because if you look if you you take this point in isolation you know there's no way that's going to be significant because it's it's right in the middle of you know in this in this thing so so this is sort of an eye-opening experiment and it's causing us to go back and re-analyze uh using this methodology it's called gene set enrichment analysis some old data sets let me give you a couple of other examples of unpublished examples that are in slightly different use it in a slightly different way so i told you about our medulloblastoma outcome prediction experiments before and around the same time there was a paper published that looked at us the same question essentially non-metastatic versus metastatic medulloblastoma different patients different arrays different group whatever they made a classifier that was centered around the pdgf receptor alpha gene was a predictor and also a number of the downstream players of pdgf receptor alpha and so we asked are any of when we look at our classifier of outcome which i showed you is pretty decent um where's the pdgf receptor alpha pathway um on there and neither pgf receptor alpha or the genes in that pathway were among the top predictors like top 50 genes in our data set which i think would lead one to believe that one or both of those data sets is wrong or the models derive from them are wrong but if you take this pdgf receptor alpha related genes as a gene set and ask is it enriched in our data set using this methodology shown schematically here it's enriched see they're not at the very top this list is is you know 12 000 genes long or so so you can see it's they're not all stacked up like 1 through 50 but they're non-randomly distributed which uh we take to mean that actually the two data sets are consistent if you had data sets of infinite size then you'd start to see convergence of the markers being overlapping at the very top of the list but with these smaller data sets and the clinical variability this puts some uh formalism around what boston was waving his hands and yelling about for so here's another example of that sort of thing is also unpublished so we um looked at lung cancer human adenocarcinoma the lung and identify some let's we just drew the line at 50 because a nice number predictors of outcome in the boston lung cancer patients university of michigan did the same experiment published around the same time overlapping gene of these two lists of 50 genes zero concerning but if you look in the space of gene set space and you ask what gene sets are enriched in one data set what gene sets which you can think of loosely as pathways they're not really pathways but it's reasonable think of them for this purpose there's really quite significant overlap in gene set space so i think what this is saying is that you know botstein is right that there is more biologic coherence in these data sets it's just we haven't been smart enough to really know how to see it so this comes from we now have about 450 or so uh such gene set some of which are good some of which aren't particularly useful they include some go annotation i don't think those are particularly useful because the granularity isn't fine enough i think in the end the most useful types of gene sets are going to be those that are experimentally derived but this is a mixture of those and we're not yet at the point where we even started to try to understand like of these 35 enriched sets of genes what are they what's the biological story so if i read it like you threw at it on the order of 60 gene sets no no no um we threw at it 400 and something gene sets and asked how many of those are enriched in the boston data set and the answer is 35 plus 18 and uh 35 plus 12 were enriched here and so the majority of the sets are enriched in one were also enriched in the other that's helpful no they're not cancer specific [Music] they're a combination as a set of you know some are like these metacarta pathways that are sort of so-so annotation some are entirely computationally derived that is they're the nearest neighbor genes of a given index gene uh in a data set they are um various things and you know what the definitive collection of gene sets would actually look like isn't obviously [Music] that's right and so i think the way to do this i mean if on the one hand i i think it will be useful to just not fret about it too much and worry about uh exactly how to define these things just get them in there um the nice thing about this gsta methodology that i didn't really go through in detail is that it's forgiving the the how you calculate these enrichment scores is forgiving of the definition of the gene sets because you're looking for enrich a non-random enrichment of the gene set so the fact that you know a third or a half of the gene set may actually be inappropriately there doesn't make any difference because there's still enough that's significantly rich to detect it you can we didn't happen to do it here but you can i'm you know again like any of these other things they're going to be a number of different metrics that you could apply to measure significant enrichment the most important thing is just to make sure that you correct for the possibility of whatever metric you use that you're detecting something beyond what you'd expect by chance so um i have two questions so i'm assuming that you can also detect like you have a location of a particular gene which you haven't mentioned so far as a given example of you know actually this is where um is that something you would find by going back to like uh your health examples and comparing you know looking for an enrichment relative to your disease yeah so actually the way you calculate this is the metric doesn't specifically look for enrichment towards the top it looks for a non-random distribution so you would find you would find um depletion we can also find which i think is this score will capture and is not desirable would be like something that's yeah concentrated in the middle which is very uninteresting so there are some false positives in be so it in fact for example in t cell activation you get you know which is you know like monolithic circulation you get up regulation of certain proteins you have down regulation of others and so what what i haven't heard yet is how you account for perhaps um you know enrichment or part of your juice yeah so there is another version of this that tries to dissect the gene sets into those components that move coherently in one direction versus the other because you're absolutely right if for example yeah if for example you take a use go annotation or something like that or some pathway if half the genes in the pathway go up and half the genes go down that could look like no enrichment at all whereas if you separate those somehow you could see it how are we doing for ton where you've got 20 minutes eight minutes okay um so let me push this to to uh not classification but some newer directions that were thinking about how can you use these signatures for useful things um particularly to think about something that's sort of closer to drug discovery so this is the way the usual discovery pipeline would look like you have some disease process or biological process you care about you do some microarray experiments and then a miracle supposed to occur whereby you develop sufficient molecular understanding of what the data are telling you that you can identify the uh smoking gun target and then you know you partner with a drug company and say screen for a small molecule that inhibits this critical therapeutic target the problem is that this part is is really tough and so what we've been thinking about is well could you bypass the understanding part at least initially um whereby you screen for small molecules based on their ability simply to perturb a signature of interest and then once you have those in hand you could use them to further dissect the biology or if you're lucky you know think about them like drugs and so the proof of concept experiment is shown here where here are two biological states for example a leukemia cell which is undifferentiated and a normal blood cell which is fully mature and has differentiated along the myeloid pathway a peripheral blood neutrophil it's not known what the critical targets are of this pathway so it's hard to do a small molecule screen to induce this process which would be nice if you could simply induce your leukemia cells to turn into normal cells so the question is could we define a signature of this state the signature of this state and then screen for compounds that trigger the signature the details don't matter here but the concept is define signatures of the two states of interest so we call this thing gehts for gene expression based high throughput screening define the signatures stands now standards what we've been talking about on microarrays where the experiment would be treat cells with various different chemical compounds and ask whether any of those compounds trigger the signature of interest and to make this feasible we simplify these complex signatures into a handful of genes that you can measure by multiplexed pcr so what's that i mean i read the paper was this purely a cost issue as opposed to going directly for microarrays it's a cost and throughput issue yeah so if you wanted to be good very completely uh if you wanted to screen tens of thousands of compounds not really feasible if it costs you you know 500 bucks a pop in the throughput issue there so yeah it's a practical matter here this part doesn't matter it's nice to say there's a method for how to measure simplified signature in high throughput so we screened a couple thousand compounds and asked do any of them trigger this little mini 5 gene signature and some of them did details don't matter then the question should be well maybe these things just trigger these five genes that are i mean sorry that these compounds trigger the five genes but actually don't do anything so one way that you can start out when they actually do anything biologically is to now step back and look across the whole genome again take cells treat them with these candidate compounds and ask did you actually recapitulate the overall molecular program not just of these five genes but of the whole thing of the whole molecular program so if you turn back to genome-wide arrays you can see that a number of these compounds recapitulated the molecular program of differentiation does that make sense so you use the simplified high thro

Original Description

MIT HST.512 Genomic Medicine, Spring 2004 Instructor: Dr. Todd Golub View the complete course: https://ocw.mit.edu/courses/hst-512-genomic-medicine-spring-2004/ YouTube Playlist: https://www.youtube.com/watch?v=_-gQchCLmXk&list=PLUl4u3cNGP613PJMNmRjAIdBr76goU1V5 I'm going to start by giving examples of a couple of patients that I saw in the Jimmy Fund clinic at the Dana-Farber that were typical. License: Creative Commons BY-NC-SA More information at https://ocw.mit.edu/terms More courses at https://ocw.mit.edu Support OCW at http://ow.ly/a1If50zVRlQ We encourage constructive comments and discussion on OCW’s YouTube and other social media channels. Personal attacks, hate speech, trolling, and inappropriate comments are not allowed and may be removed. More details at https://ocw.mit.edu/comments.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from MIT OpenCourseWare · MIT OpenCourseWare · 0 of 60

← Previous Next →
1 21. Post Trade Clearing, Settlement & Processing
21. Post Trade Clearing, Settlement & Processing
MIT OpenCourseWare
2 10. Financial System Challenges & Opportunities
10. Financial System Challenges & Opportunities
MIT OpenCourseWare
3 7. Technical Challenges
7. Technical Challenges
MIT OpenCourseWare
4 3. Blockchain Basics & Cryptography
3. Blockchain Basics & Cryptography
MIT OpenCourseWare
5 19. Primary Markets, ICOs & Venture Capital, Part 1
19. Primary Markets, ICOs & Venture Capital, Part 1
MIT OpenCourseWare
6 1. Introduction for 15.S12 Blockchain and Money, Fall 2018
1. Introduction for 15.S12 Blockchain and Money, Fall 2018
MIT OpenCourseWare
7 Chalk Radio, A Podcast about Inspired Teaching at MIT (Teaser)
Chalk Radio, A Podcast about Inspired Teaching at MIT (Teaser)
MIT OpenCourseWare
8 Nuclear Gets Personal with Prof. Michael Short (S1:E1)
Nuclear Gets Personal with Prof. Michael Short (S1:E1)
MIT OpenCourseWare
9 How Africa Has Been Made to Mean with Prof. Amah Edoh (S1:E2)
How Africa Has Been Made to Mean with Prof. Amah Edoh (S1:E2)
MIT OpenCourseWare
10 Making Deep Learning Human with Prof. Gilbert Strang (S1:E3)
Making Deep Learning Human with Prof. Gilbert Strang (S1:E3)
MIT OpenCourseWare
11 Social Impact at Scale, One Project at a Time with Dr. Anjali Sastry (S1:E4)
Social Impact at Scale, One Project at a Time with Dr. Anjali Sastry (S1:E4)
MIT OpenCourseWare
12 Film is for Everyone with Prof. David Thorburn (S1:E5)
Film is for Everyone with Prof. David Thorburn (S1:E5)
MIT OpenCourseWare
13 Lecture 12: Aircraft Performance
Lecture 12: Aircraft Performance
MIT OpenCourseWare
14 Lecture 3: Learning to Fly
Lecture 3: Learning to Fly
MIT OpenCourseWare
15 Lecture 13:  Interpreting Weather Data
Lecture 13: Interpreting Weather Data
MIT OpenCourseWare
16 Lecture 21: Weather Minimums and Final Tips
Lecture 21: Weather Minimums and Final Tips
MIT OpenCourseWare
17 Hand-on, Minds On with Dr. Christopher Terman (S1:E6)
Hand-on, Minds On with Dr. Christopher Terman (S1:E6)
MIT OpenCourseWare
18 Part 4: Eigenvalues and Eigenvectors
Part 4: Eigenvalues and Eigenvectors
MIT OpenCourseWare
19 Part 5: Singular Values and Singular Vectors
Part 5: Singular Values and Singular Vectors
MIT OpenCourseWare
20 Part 3: Orthogonal Vectors
Part 3: Orthogonal Vectors
MIT OpenCourseWare
21 Part 2: The Big Picture of Linear Algebra
Part 2: The Big Picture of Linear Algebra
MIT OpenCourseWare
22 Part 1: The Column Space of a Matrix
Part 1: The Column Space of a Matrix
MIT OpenCourseWare
23 Intro: A New Way to Start Linear Algebra
Intro: A New Way to Start Linear Algebra
MIT OpenCourseWare
24 9. Chromatin Remodeling and Splicing
9. Chromatin Remodeling and Splicing
MIT OpenCourseWare
25 28. Visualizing Life - Fluorescent Proteins
28. Visualizing Life - Fluorescent Proteins
MIT OpenCourseWare
26 20. Roth's theorem III: polynomial method and arithmetic regularity
20. Roth's theorem III: polynomial method and arithmetic regularity
MIT OpenCourseWare
27 8. Szemerédi's graph regularity lemma III: further applications
8. Szemerédi's graph regularity lemma III: further applications
MIT OpenCourseWare
28 19. Roth's theorem II: Fourier analytic proof in the integers
19. Roth's theorem II: Fourier analytic proof in the integers
MIT OpenCourseWare
29 12. Pseudorandom graphs II: second eigenvalue
12. Pseudorandom graphs II: second eigenvalue
MIT OpenCourseWare
30 1. A bridge between graph theory and additive combinatorics
1. A bridge between graph theory and additive combinatorics
MIT OpenCourseWare
31 Special Episode: Teaching Remotely During Covid-19 with Prof. Justin Reich
Special Episode: Teaching Remotely During Covid-19 with Prof. Justin Reich
MIT OpenCourseWare
32 Spring 2020 Update from Dean Rajagopal
Spring 2020 Update from Dean Rajagopal
MIT OpenCourseWare
33 S1E7: Unpacking Misconceptions about Language & Identities with Prof. Michel DeGraff
S1E7: Unpacking Misconceptions about Language & Identities with Prof. Michel DeGraff
MIT OpenCourseWare
34 Climate 101 Live
Climate 101 Live
MIT OpenCourseWare
35 Welcome for Volunteers (for EarthDNA's Climate 101)
Welcome for Volunteers (for EarthDNA's Climate 101)
MIT OpenCourseWare
36 Learning to Fly with Drs. Philip Greenspun & Tina Srivastava (S1:E8)
Learning to Fly with Drs. Philip Greenspun & Tina Srivastava (S1:E8)
MIT OpenCourseWare
37 Thinking Like an Economist with Prof. Jonathan Gruber (S1:E9)
Thinking Like an Economist with Prof. Jonathan Gruber (S1:E9)
MIT OpenCourseWare
38 2. Cyber Network Data Processing; AI Data Architecture
2. Cyber Network Data Processing; AI Data Architecture
MIT OpenCourseWare
39 1. Artificial Intelligence and Machine Learning
1. Artificial Intelligence and Machine Learning
MIT OpenCourseWare
40 2: Resistor Capacitor Circuit and Nernst Potential - Intro to Neural Computation
2: Resistor Capacitor Circuit and Nernst Potential - Intro to Neural Computation
MIT OpenCourseWare
41 14: Rate Models and Perceptrons - Intro to Neural Computation
14: Rate Models and Perceptrons - Intro to Neural Computation
MIT OpenCourseWare
42 4: Hodgkin-Huxley Model Part 1 - Intro to Neural Computation
4: Hodgkin-Huxley Model Part 1 - Intro to Neural Computation
MIT OpenCourseWare
43 18: Recurrent Networks - Intro to Neural Computation
18: Recurrent Networks - Intro to Neural Computation
MIT OpenCourseWare
44 3: Resistor Capacitor Neuron Model - Intro to Neural Computation
3: Resistor Capacitor Neuron Model - Intro to Neural Computation
MIT OpenCourseWare
45 15: Matrix Operations - Intro to Neural Computation
15: Matrix Operations - Intro to Neural Computation
MIT OpenCourseWare
46 13: Spectral Analysis Part 3 - Intro to Neural Computation
13: Spectral Analysis Part 3 - Intro to Neural Computation
MIT OpenCourseWare
47 16: Basis Sets - Intro to Neural Computation
16: Basis Sets - Intro to Neural Computation
MIT OpenCourseWare
48 20: Hopfield Networks - Intro to Neural Computation
20: Hopfield Networks - Intro to Neural Computation
MIT OpenCourseWare
49 8: Spike Trains - Intro to Neural Computation
8: Spike Trains - Intro to Neural Computation
MIT OpenCourseWare
50 7: Synapses - Intro to Neural Computation
7: Synapses - Intro to Neural Computation
MIT OpenCourseWare
51 19: Neural Integrators - Intro to Neural Computation
19: Neural Integrators - Intro to Neural Computation
MIT OpenCourseWare
52 5: Hodgkin-Huxley Model Part 2 - Intro to Neural Computation
5: Hodgkin-Huxley Model Part 2 - Intro to Neural Computation
MIT OpenCourseWare
53 6: Dendrites - Intro to Neural Computation
6: Dendrites - Intro to Neural Computation
MIT OpenCourseWare
54 17: Principal Components Analysis_ - Intro to Neural Computation
17: Principal Components Analysis_ - Intro to Neural Computation
MIT OpenCourseWare
55 12: Spectral Analysis Part 2 - Intro to Neural Computation
12: Spectral Analysis Part 2 - Intro to Neural Computation
MIT OpenCourseWare
56 11: Spectral Analysis Part 1 - Intro to Neural Computation
11: Spectral Analysis Part 1 - Intro to Neural Computation
MIT OpenCourseWare
57 9: Receptive Fields - Intro to Neural Computation
9: Receptive Fields - Intro to Neural Computation
MIT OpenCourseWare
58 10: Time Series - Intro to Neural Computation
10: Time Series - Intro to Neural Computation
MIT OpenCourseWare
59 1: Course Overview and Ionic Currents - Intro to Neural Computation
1: Course Overview and Ionic Currents - Intro to Neural Computation
MIT OpenCourseWare
60 The Power of OER with Profs. Mary Rowe and Elizabeth Siler (S1:E10)
The Power of OER with Profs. Mary Rowe and Elizabeth Siler (S1:E10)
MIT OpenCourseWare

This video lecture teaches how to apply genomic medicine in cancer diagnostics using microarrays, gene expression analysis, and RAG search. It covers the use of these tools to understand the molecular determinants of cancer outcome and develop targeted therapies.

Key Takeaways
  1. Cluster patients into two groups using a self-organizing map
  2. Apply supervised learning with a k-nearest neighbor classifier to predict patient survival
  3. Use an 8-gene model with cross-validation to classify patients
  4. Calculate statistical significance using the log rank test on a Kaplan-Meier survival curve
  5. Rank genes by correlation with distinction and interrogate gene sets for enrichment
💡 The use of RAG search and gene expression analysis can provide valuable insights into the molecular determinants of cancer outcome and help develop targeted therapies.

Related Reads

📰
What Is RAG AI? Retrieval-Augmented Generation Explained — American Dream AI
Learn about RAG AI, a technology that enhances AI generation with retrieval capabilities to improve accuracy and reduce hallucinations
Medium · AI
📰
Treat Retrieved Content as Data, Not Instructions
Learn to treat retrieved content as data, not instructions, to improve model inference and context assembly
Dev.to AI
📰
How to Evaluate Production RAG: Keyword, Vector, SQL, and Hybrid Retrieval
Learn to evaluate production RAG systems by testing keyword, vector, SQL, and hybrid retrieval routes against the same questions
Dev.to · Anya Summers
📰
RAG Database Design: SQL, Full-Text Search, Vector Search, and Context Retrieval
Learn to design a RAG database with SQL, full-text search, vector search, and context retrieval for efficient information retrieval
Dev.to · puffball1567
Up next
Build a Chatbot with RAG in 10 minutes | Python, LangChain, OpenAI
Thomas Janssen
Watch →