14: Motion Perception (cont'd)
Key Takeaways
The video discusses motion perception, covering topics such as the Reichart detector, motion detectors, and the aperture problem, as well as the role of form information and contour completion in motion integration, using concepts from retrieval augmented generation and fine-tuning.
Full Transcript
So today we're going to um finish talking about motion perception. Um so we got this started last time. Uh so motion happens when things in the world change position over time. Um and the problem of motion perception is understanding how we see that motion. Okay. Um and so we talked about like the the evidence that there are motion detectors in the visual system in the way of the motion after effect. We experienced a motion after effect. Um and we we talked about simple models for how you would construct such motion detectors. Sort of the simplest and most intuitive is what's called the Reichart detector where there are spatially displaced inputs that arrive at a downstream neuron uh with different time delays such that if there's something that's moving through space such that the spatial offset kind of matches the temporal offset um the inputs end up reaching the downstream neuron at the same time and if the downstream neuron is set up as a coincidence detector then you'll get a a direction selector response. Okay. Um, we also looked at an actual direction selective neuron and the receptive field that you can see if you you measure the spike triggered average of such a neuron and that consists of in V1 in this example of these orientation selective receptive fields that kind of shift over time. Okay. Um, and so then we talked about the idea that like another way to think about motion detection is that you can think of motion as being orientation in spaceime. Um, and so to detect motion, you really need filters that are oriented in spaceime. Okay, so that's another way to kind of think about this. And these actual receptive fields that are measured in actual neurons um, show this orientation in spaceime. So remember, one of these receptive fields is like a movie, right? So you got the x and the y dimensions and the t dimensions. So it's like a bunch of of image frames. Uh, and but you can project those down to one spatial dimension just so it's easy to look at. And then you see this orientation in this case in X and T. All right. So then we talked about the so so the this this is evidence that these mechanisms in your in primary visual cortex for detecting image motion um get oriented inputs, right? So you could kind of imagine kind of building this um from a bunch of simple cell inputs and the consequence of them being oriented um is that there's an ambiguity in the motion uh that they signal, right? All right. And so specifically, we talked about the idea that one a motion detector that's set up like this. It can tell you how much something is moving in this case in the horizontal direction. All right? It's not going to tell you how much it's moving in the vertical direction. Okay? So they're sort of measuring a component of the velocity rather than the velocity in its entirety. Okay? And so the consequence of that is that if you observe a response in one of these local motion detectors um that response is giving you a constraint line in the space of velocities. Right? So if you think of you know velocity is a vector quantity. Okay? So it's got two dimensions. There's an x component and a y component. Okay? Um and the constraint line means that the the data that you observe. So the the response that you observe is consistent with each of these possible vectors. Right? All right. And what do these bat vectors share? They share the horizontal component in this particular example. Um but the vertical component is unconstrained. Okay. All right. So each one of these these local motion measurements is giving you information about motion, but it's not uniquely specifying the 2D motion. And then we talked about this idea that if you have multiple detectors um that are tuned to different orientations, they will give you distinct constraint lines. And so if both of those detectors are stimulated by the same thing, then you can take the intersection of those constraint lines and that would tell you the two-dimensional direction of motion that's happening in the world. Okay? So you can think about this in terms of neurons um as taking a whole bunch of different um local motion detectors each one of which is oriented and kind of adding them all up adding adding all of the ones up that are consistent with a particular 2D motion direction. So each one of these things has a particular constraint line and so this particular combination here is all the ones that intersect at this particular point. Okay. So you could build something that is selective for that particular 2D velocity with this particular combination of these simple motion detectors. And if you wanted something that was tuned to some other velocity like up here, that would give you a different combination. Okay, so this is kind of a theoretical idea suggesting that there are these two stages of of motion processing. an initial stage where things are decomposed into one-dimensional components and then a second stage um when those are combined potentially using this intersection of constraints idea that we talked about. All right. And so then we began to talk about the um the evidence that this these two stages that are sort of postulated on theoretical grounds um actually have a physiological basis in the visual system. Um and the the story here um is in this area called MT which is a visual area that's kind of midway up the dorsal pathway. Remember we think of the visual system as being organized into these two pathways. Um the vententral pathway which we think mediates object recognition. The dorsal pathway also often called the wear pathway but it's often um it's also where motion analysis seems to take place. And so this is area MT is part of the dorsal pathway. And so we saw how in area MT virtually all of the neurons are selective for the direction of motion. So they're tuned to direction. Okay, so that's one indication that it probably is important in the perception of motion. Uh and I didn't show you direct evidence for this, but I asserted that if you leion MT, uh you get deficits in in motion perception. Okay. Um and so there were when we left off last time was by introducing the idea that you could study this problem using this stimulus which is known as a plaid. Okay. So a plaid is a stimulus that is composed of two sine wave gradings. So here's one sine wave grading and here's another. And you get the plaid by just adding them together. Okay. And so the idea behind the stimulus is that each of the syosoidal components will stimulate a different elementary motion detector. Um and then uh when the gradings are superimposed um the question is whether or not the the the outputs of those two motion detectors gets combined and if so where in the visual system. Okay. Um so this is what each of those looks like on its own and then when you add them together the interesting thing is that you see a single pattern um that moves in a direction that's different from the direction that you see with either of the components on its own. Okay. Yeah. are a little bit more like for example different colors and I guess what might explain those difference. >> Yeah, that's a that's a good question. So the the extent to which um the the two gradings cohhere so look like a single pattern will depend on the property of those gradings. So, so for instance, if the spatial frequency of the gradings is more different, so if one of them is very low and one of them's very high, you'll be less likely to see the pattern direction of motion. Um, I don't actually know whether differences in color will have an effect. Um, but other things that you might think would be related to um the extent to which those belong together uh do affect this. Okay. So, it's not like there's some obligatory integration and it's affected by principles of perceptual organization. Yeah. Okay. So, yeah. >> Um, plaid motion. Yeah. Yeah. This is a it's a very famous stimulus in sort of the history of vision science. Okay. So, the point is that like this is a stimulus. Perceptually, you see it as moving in this direction of the pattern. Okay. It but it's composed of these two components. All right. And the idea is that those components are kind of what like elementary motion detectors are going to see, right? Because they're tuned to orientation. Okay. And so the question is, is there a place in the visual system where the outputs of these elementary motion detectors would get combined and respond to the pattern direction? Okay. So just to sort of confirm the the intuition or to illustrate the intuition a little bit more. So we've got this stimulus here, right? that consists of these two components. Um, and it it creates a plaid. Okay. And so if you have a a neuron that is selective for the direction of motion, you might for instance see tuning like this. So this is a tuning function in directions. This is polar coordinates. Okay? So the direction here is indicated by the angle. And then this is a plot of the response of the neuron at each direction of motion. And so the fact that this is way out here for this direction is like an indication this neuron is tuned for horizontal motion that is to the right. Okay. So that's a a directional tuning curve. And this is what you would measure with a single grading. That's why it's called the grading response. Okay. And so there are two kind of possibilities for what you might get if you show this a plaid show you show this neuron a plaid that's composed of these two gradings. Okay. So one possibility is that the neuron is going to respond um to each of the gradings individually in which case there will be this billobed response function because what happens is that you rotate the plaid around and you move it in every possible direction and then there will be two directions in which one of the gradings is aligned to the the tuning of the of the neuron. Okay, so this is what you would expect like say in area V1, right? that there's like two places where two directions where you get a big response where one of the gradings is aligned with the preferred direction or where the other one is. Okay. Um but the other possibility is that the neuron might be responsive to the plaid direction, right? So the direction that you actually see when you look at this and if that was the case then you would get a single lobed response um that would be aligned to the direction of the tuning that you would get with a single grading. Okay. All right. Right. So those are sort of the two idealized possibilities. And the question is like what do you see? Okay. So in area V1 um you you almost exclusively see neurons that are seem to be responsive to the individual gradings of the plaid. Okay. Um so in the way this this is evaluated with that same experiment. So you measure the directional tuning for with a single grading. So this is a neuron that prefers this kind of direction that's downward and to the right. Um and then you measure the direction tuning with plaids. And the directional tuning that you get with plaids has these two loes. Okay. Um and one of these is the actual response. I think the solid line is the response and the dash line is kind of the prediction um of the response just from the response to gradings under the assumption that it just responds to the gradings. Okay. So again, the idea is that these neurons in V1, they're they're orientation selective filters. And so they essentially just see one orientation component at a time when one of them aligns with the receptive field. Okay. And so this can be quantified um by measuring the correlation of the observed response function um with the with with these two predictions here, right? With the prediction of what you would what you would expect if the neuron were just responding to the components and by contrast the prediction if it was just responding to the plaid. Okay? Okay, so you get the correlation between the actual tuning curve and these two predictions and um you can plot that in this plane here. So this is the correlation with the component prediction. This is the correlation with the pattern prediction. Okay. And so the idea is that if this is a neuron that is really just responding to the components, then it's going to be in this region of space because it should be highly correlated with the component prediction. Whereas if it's responding to the pattern direction, it should be kind of up here. Okay. And so this is a plot of the results of doing this analysis on neurons in area V1. So each dot here is a neuron. Okay. And what you can see is that pretty much all of the dots are in this region here. And what that means is that the directional tuning curves of V1 neurons to plaids have these by this blob shape that resembles the prediction if they were just responding to the components. Okay. So the point is that like V1 is doing what you would expect and and not doing something s um uh interesting. Now the question is like what happens downstream. Okay. And so this is this area MT um and so that some of the neurons in MT look just like neurons in V1. So here's an example. So this is the tuning curve to gradings and this is the tuning to plaids and you get again this blobed tuning response function. Okay. Um, but here's a case where the direction tuning to plaids looks a lot like the direction tuning to gradings and in fact deviates quite significantly from what would be predicted if it was just responding to the components. Okay, so this is an example of a neuron that is tuned to the direction of motion that you kind of see when you look at that pattern. Okay, so this is just two examples. Um this is the result of this analysis where um every dot again is a neuron. Okay. Um and although you can see some of the the neurons here are in this this lower region indicating that they're responsive to the individual components just like in area v1 there's a bunch that are up here in this top region indicating that they're responsive to the pattern direction. Okay. Okay. So this is evidence for this two-stage model of motion perception, right? There's this initial stage where you measure responses to 1D components. Um those then seem to provide inputs to the second stage potentially located in area MT uh where those responses get combined and you end up with neurons that are responsive to the pattern direction. Okay. Okay. So as I as I mentioned last last time, this is was a fairly influential story in the the history of sort of systems neuroscience, sensory neuroscience um in that it was a nice example where there was convergence between both perception um so this stimulus of plaids was introduced and show shown to to exhibit these interesting properties, these theoretical ideas for how you would measure motion. Um, and then some experimental evidence where different stages of the brain kind of mapped onto these different theoretically motivated computational stages in like a pretty clear way. Okay. So, sort of a a classic success story in in u systems neuroscience. Any any questions about how this works? >> Yeah. So if the if the plat motion response neurons are only seeing like half of the grading, will they respond partially or not at all? >> Um, by half of the grading, you mean by only one grading? >> Yeah, only one grading. >> Well, that's sort of that's what's shown here. [snorts] So this first column is the tuning function if they're only presented with one grading. Okay. Um and so they respond to the to the single grading, right? The key thing that that um they do is that the direction tuning to a single grading is the same as the direction tuning to the plaid. Even though when the plaid is moving in this direction, right, which is the direction that it responds a lot to when it gets a grading, the individual gradings of the plaid are one of them is moving this direction and one of them is moving in this direction, right? So that's what's kind of remarkable about it, right? Is that it's doing this kind of integrative operation um that's nonlinear uh and uh that gives you this kind of emergent property. Okay. All right. So, story so far um is that we've got this two-stage process for detecting the direction of motion. There's this area called MT where neurons are able to detect 2D motion. Um and what we're going to talk about now is this additional challenge that um is present, right? And so, previously we've been sort of talking about this challenge that happened kind of in some sense because of the way the visual system was set up, right? So you start out making these motion measurements that are oriented and then that has this consequence that you have to kind of combine across orientations in order to get um 2D motion signals. So there's this other problem though um which is that there's an ambiguity that exists in the world. Okay. Um and so so the neur the issue here is that neurons typically are making local measurements. Right? We talked about this idea that like the neuron looks at the world through a receptive field that is sort of restricted in area. All right. And on your homework uh on your problem sets, you've sort of done this exercise of looking at images through apertures and verified like how things are frequently pretty ambiguous when you just have local evidence. Okay. Um and so another type of this ambiguity um is due to the fact that many things in the world in particular the edges of things are locally one-dimensional. Um and um as a consequence there is they're inherently ambiguous. um their their motion is inherently ambiguous independent of how you measure it. Okay? And this is known as the aperture problem. So here here's the idea. So we have an edge here at time one. Okay? And the edge moves. All right? So this is where the edge is in the image at time two. But if you're viewing it through this aperture which could be like a receptive field again like all that you can measure in principle is the component of the of the velocity that is perpendicular to the orientation. Okay. Um and just due to the geometry here you don't really know what the component is that's parallel to the orientation. Okay. And so the consequence is that the edge could be moving in any of these particular directions. Okay. All right. This is called the aperture problem. [snorts] So, here's just a an example of this. Um, so these three motions, they're all different, um, but they look the same when you view them through the aperture. And I can show you some actual examples of this that involve motion. Okay. So, oh no, this happened before. Do you remember this from last year? What? What >> a flash drive plugin. >> Yeah, I know this is it, but it's it's just that it's um Oh, there we go. Okay, so um this is a demo of uh a simple scene with some moving objects um that you can view through different apertures. Okay, so this is a one aperture. You're looking at this thing and it kind of looks like it's moving like diagonally, right? Um and where's my there's my mouse. Okay. Um so this is what is actually there. Okay. So everything here is moving horizontally. Um but the edges are ambiguous. Now not everything is ambiguous locally, right? So there are certain features that are two-dimensional which are unambiguous. Um but some things are highly ambiguous and that's the aperture problem. Okay. So um given this issue right that that edge motions are ambiguous how does the visual system determine their velocity? And so there are two main answers that we'll talk about. Um the first is to make use of unambiguous 2D signals. So like the corner that you saw there to generate an unambiguous solution. And the second one is to combine edge motions across space. So if you have different edges that are different orientations, you can combine those and and resolve the ambiguity. Okay? So each edge is ambiguous on its own, but together they uniquely determine a velocity. Um so this is a really pretty cool example of a real life consequence of the aperture problem and our reliance on 2D features. So this is um I think this is at like a a football stadium in Spain or something some something in Europe. Okay. Okay. And the these are exit ramps to get out of the stadium. And so they're just these big long ramps that are spirals. And so people just walk down the ramps. Okay. Um but because of the aperture problem like the mo the you don't the motion of these edges is ambiguous. Um and it gets captured by the motion of the people who are walking. Okay. And so the whole thing looks like it's spiraling. So what's actually happening there is that you know this is this huge gazillion ton structure in the world that of course is not moving the people are just walking right but because of the smoothness of the edge and the aperture problem um the 2D motions of the people sort of capture the whole thing and your visual system ends up thinking the thing is is spiraling down into the ground or something right so yeah quite amazing okay now one of So, so we we often rely on these 2D signals. Now, another challenge that comes up as a consequence of this is that not all of the 2D motion signals that are present in images are real motion signals in the sense that they correspond to like an object that's actually moving in that direction. Um, and one of the the main causes of this is that uh you can get lots of 2D motion signals from occlusion. Okay, so this is the same situation where we have two squares. they're moving horizontally, but the intersection, the place um where where they where one occludes another will actually be moving up or down. Okay? And so you can see this in that demo that I was just trying to show you. And okay, so um [clears throat] yeah, so this thing looks like something that's moving up and down, right? But it's actually just this, right? Okay. Um, so anytime things olude one another and move that this sort of tends to happen. Okay. Um, and so you you know, you might expect that, you know, in order to see things moving correctly then the process of determining or inferring motion would need to be sensitive to occlusion, right? The fact that sometimes one object is in front of another. Um, and so this is, and there's now lots of demonstrations that this is the case. This is one that's quite powerful where I'm going to show you a stimulus that consists of these two bars, one that's moving vertically, one that's moving horizontally. Um, and you can interpret when when you see both of the bars at the same time, you can interpret the motion in one of two different ways. Um, depending on whether there's evidence for occlusion. Okay? So, if you see the bars on their own, they look like they're moving separately. One's up and down and one's back and forth. But then when we put this little frame around them which plausibly could be oluding the end points um you now see the thing as one object that's kind of moving in a circle. All right. And so the plausible explanation is that your visual system supposes that the image motion that's happening here is due to occlusion and that no longer really has a big impact on what you see. Okay. So this is um an example of that. And I should just say that um these demos hang on a second. There we go. Okay. So this is the basic stimulus. And so this is what happens if the bars are just on their own. Hopefully everybody just agrees. It looks like two different things, right? um we add this frame around them and most of you probably see like a single thing that kind of moves around in a circle. All right. So the thing to emphasize here is that the image motion is exactly the same in these two cases. Okay. Um but you arrive at these two different interpretations um depending on whether there's evidence for occlusion or not. Okay. I just wanted to I just wanted to mention these demos. These were created by an MIT Europe in like 2002 or something. Okay. Um and uh they uh there used to be a website that had all of these things and that but it was their program in Flash and Flash doesn't work in browsers anymore. So you can download a Flash player and they still work. But um this is just to say that you never know the things that you do in your Europe might still remain relevant in 20 years. So kind of cool. Okay. So the point of this is that motion analysis seems to be informed by information about depth and occlusion. Motions that occur that where you have evidence that they occur at points of occlusion tend to be discounted. So if we go back to the outline that I was just showing you, we talked about that this idea that the aperture problem can kind of get resolved with kind of two in two main ways. One is to make use of unambiguous 2D signals. Um we saw some examples of that. The trick there is that you need to discount 2D signals that come from occlusion, which means you have to take into account information about occlusion. Um, but then the other way you could overcome the aperture problem is by integrating edge motions across space. Okay. Um, and so the challenge there um is that sometimes local motions arise from different objects. So if you just take this same display again with these two squares that are moving in opposite directions, here's the situation, right? So you can make let's suppose you have a receptive field here and a receptive field here and a receptive field here. Right? So the consequence of the aperture problem is that each one of these measurements is ambiguous. Um it's not useless not uninformative but it's ambiguous. And so what that means is it gives you a constraint line. Okay. So with this measurement you get this particular constraint line. So the the this lo this local motion is consistent with any velocity anywhere on this line. Okay. and this particular local measurement is consistent with a different constraint line. All right, so the intersection of those uh is unambiguous and gives you the correct direction of motion which is this horizontal motion. Okay, so life is good. Okay, but let's just suppose that um you instead want to decide to combine the measurement that you make here with the measurement that you make here. So this one again gives you that same constraint line here, but this one gives you a different constraint line. And now the intersection of the constraints is vertical, right? And nothing in the scene is moving vertically. Okay. All right. So this business of integrating information really only kind of makes sense um if if you kind of know what to combine with what. Okay. Um and one possibility here is that you would make use of form information to deter to determine that. There's now lots of um cool demonstrations of this. This is a classic one. This was introduced by Maggie Shiffar and Joano. So again, there these moving bars. Um and this is a situation where um when there is evidence for occlusion um that could potentially account for um the fact that there would be a single diamond here, right? That just the corners of which happen to be covered. Um you tend to see these things as moving as a single thing. Um and without that evidence for occlusion, you tend to see separate things. Um, and that's a super powerful effect. So, here's how that works. [snorts] Okay, so two sets of bars, they look like they're moving independently. Then we introduce occlusion and suddenly you can see like a single thing that's kind of moving around in a circle. Okay. Um um importantly, you know, there's really nothing there's certainly nothing locally that moves in a circle, right? So that circular motion that you perceive has to result from combining information across the edges and you combine it um in this setting but not in this setting. Okay? All right. And so intuitively like one possibility is that this process of integrating the motion of the edges might be linked to contour completion. So a couple lectures ago we sort of talked about the fact that when things are oluded there's this process of aodal completion that kind of allows you to to sense the presence of one object behind another. Okay. Um and you might think that that is happening here and that the integration of the motion is is related to that. Um and in fact there's there's a bunch of pieces of evidence for that. Um so here's sort of one variant on this stimulus that can be used to test this. So here the idea is that um there's a manipulation of the oluders so that in this case there's kind of room beneath the oluders for the there to be a full diamond right for these contours to kind of connect. Um and the idea is that here there there really isn't room. Okay. And so if you think that really the the process of integrating and and combining the motion of the edges is dependent on the contour completion, you might expect that this thing would would cohhere as a single object a lot less than that one. Okay. Um and that is in fact what tends to happen. You can verify this for yourself. Okay. So this is a case where there's kind of room for the diamond to kind of complete bes behind the oluders and most people tend to say see that say that they see a single thing kind of moving around in a circle. But if we make those oluders much thinner um that that integration probably kind of breaks. Is that I see some people nodding. Is that working for all of you? Yeah. Good. If we make them thick again, um you're probably able to sort of see a a single thing moving around in a circle. Okay. Um we make them thin, possibly breaks. Um this is another sort of manipulation. So if we make some a small change here that now allows you to kind of see the oluders as these extended surfaces again, um people tend to be able to then see a single object moving around in a circle again. Yeah. working. Yeah. Okay. Okay. And so these are these are just sort of demonstrations that you can evaluate from looking at them. But there's you know you can do experiments to measure this. These are just bar graphs that plot the proportion of time that people say that they see this what's called the coherent interpretation which just means that there's a single thing kind of moving in a circle. Um and when there's room for the thing to emotally complete behind the surfaces here or here um people tend to report the coherent interpretation and then when there isn't room they don't. Okay. Now you may remember back when we were talking about contour completion um that another thing that affects completion um is this property of relatability. So it sort of has to do with whether the contours are positioned so that they can complete. And so here the four white lines would be would be considered to be relatable because you can kind of connect them with like a smooth contour. And here they're really not relatable, right? You have to have a really kind of wacky looking shape in order for those things to all be um connected. Okay? And so if you think that the motion integration would be affected by this, you might expect that there would be a big difference in the extent to which you would see coherent motion even though the local motions in all these cases are the same, right? you're just kind of moving them around a little bit. And so here's um here's an example of that. Okay. So here's a case where things are relatable. Um how many people are are seeing this as like a single thing kind of moving in a circle? Raise your hand. Okay. Most people. In fact, we can another thing that sort of affects us a little bit is the contrast of we can even make it a bit lower contrast. So everybody kind of seeing that as moving in a circle. There's a one person who doesn't. But what about this case here? Yeah. Yeah. So that sort of causes it to break. Now you may say, okay, well that's it's just impossible to see the thing moving in a circle in that case, right? Okay. But then I would say, well, what about this? Okay, so we add these little 2D features to the thing and now everything kind of looks like it's moving in a circle. So this is a demonstration that it is in principle possible to see that as a single shape kind of moving around in a circle. Um but without that um ev that local evidence um you don't really interpret it in that way, right? We move it back to here. Yeah, much more likely to to integrate that. All right. And so there's you can do experiments that sort of verify what hopefully you just saw, right? These are these are big effects, so they're not very subtle. All right. Um so again, this is sort of evidence that these these kind of mid-level um processes of contour completion um seem to be intimately related to the interpretation of of motion. Um here's just another demonstration of of related kinds of things where uh information about about grouping in this case in depth really seems to affect things. Okay, so there I'm going to show you four stimula here. These are um each kind of a a plaid that consists of a couple simplified gradings. This is the basic plaid. And then we introduce this gray thing that can either be in front or behind or kind of in between the two gradings. Okay? Okay. And so the idea is that when you position this thing in between them, um it becomes kind of impossible for the things to be connected as part of the same thing or much less likely. Um and that ends up having a big effect on the motion interpretation. Okay. So here's this one. Okay. So hopefully most of you are perceiving horizontal motion here. Okay. So now we just add this other thing in back and you still see something horizontal. Now we put it in front. Still probably looks horizontal but now pro you probably see the two things is moving kind of separately. Yeah. Okay. Okay. So again, like you know, the the local motion information is fairly similar in a lot of these cases, but the way that it gets interpreted um is very differently sort of depending on the other kinds of evidence that these two sets of lines could be part of the same thing. Okay? And again, you can run experiments to kind of verify all of this and things kind of work out more or less like you like you would expect. Okay, so the the sort of big picture here, right? These are all kind of demonstrations um that there are these big interactions between the perception of motion and the perception of form and grouping um and shape. Um and so you know we classically think of the visual system as consisting of these two streams. Um and the machinery for processing motion kind of looks like it's in this dorsal stream. Um and naively you might expect that all the stuff for dealing with shape and so forth would be in the vententral stream. And maybe that's true in some way. Um but there's pretty clearly like some very tight linkages. Um, and you know, we don't really know how these perceptual phenomena that I just showed you are actually sort of mediated, but um, there's got to be very tight interactions. Um, and so it seems like the visual system tends to infer an explanation of motion in terms of surfaces and objects again based on implicit knowledge of the way the world works. Right. Okay. All right. Any questions about those sorts of phenomena? [snorts] All right. So um another uh thing that's worth knowing about and it's sort of thematically related to what we've just been talking about. Um so we just we just talked about these examples where form information seems to constrain the interpretation of motion. There's also very famous examples where motion alone can define form. Um and so these the original demonstrations of this came from uh experiments where this the scientists put little lights on the joints of humans. Okay. um and then film them in a dark room. So the idea was that like the only thing that was visible was the motion of these a few points on a a person's body. Okay. Um and when you do this, you can see stuff. Okay. So the first thing I'm going to show you is this, which is probably not going to look like a whole lot. Okay. Um uh this is something meaningful, but I did a I played a little trick on you. Okay. Um and this is the one that's probably recognizable. Okay. Okay, so you see this and like you immediately see that there's a person walking and the one that I showed you first is exact the exact same thing just flipped upside down. Okay. Um and so there's like a a presumably a prior that kind of favors upright walking just because that's something you see all the time. Um so that kind of constrains this. But the the key point is that you know this is again sort of illposed like there's lots and lots of possible explanations for the motions of these dots. um um but they are consistent with a person walking and they cause you to see a person walking and there's um lots of other examples of things like this. All right, so motion can define form u. Oh, this is another kind of cool example. So this is the a little movie that resulted from running a an old school edge detection algorithm on a on a movie. And um this is pretty interesting just because each individual frame of the movie is really hard to interpret. But then when you kind of see them all um as a movie, you can kind of easily recognize that you got some dogs playing. You can kind of see people in the back kind of walking along the fence. Um but if I if I like stop it in the middle, oops, I shouldn't stop it there. I pause it here. You know, it's each individual frame here is really hard to interpret. Okay. All right. So, lots of cases where where motion really helps you see shape. All right. So, what I want to turn to now though um is is a puzzle in motion perception. And the answer to that puzzle really comes in the form of a Beijian model of motion perception. Um, and so this turns out to be a pretty nice application of uh, Beijian theories of perception that we've kind of talked about like repeatedly over the course of of the class. So we talked about this idea that you can take local motion measurements and combine them according to the intersection of constraints, right? There's lots of cases where the local motion measurements gives you a constraint line. Okay? And to get the the true 2D direction of motion, you would combine the constraint lines. And so those constraint lines, those could come from different 1D measurements that are made in the visual system in V1. They could come from different 1D edges of objects in the world situations where you got the aperture problem. Um same same issue. Um and so this is uh this is a case um where what you see is actually not the intersection of constraints direction. So that's really the puzzle, right? There are these instances. So a lot so I should back up and say a lot of the time the direction that people do see like when you're looking at a plaid is given by the intersection of constraints right but then what so so historically what happened was plaids were introduced as a stimulus for studying motion this was back in the early 80s okay um and there were these initial demonstrations that the direction of motion people see when they look at a plaid is typically the intersection of constraints direction right and then everybody got excited there were like a million different experiments on plaids and people started to realize that there are conditions where sometimes what you see is actually not the direction that's implied by the intersection of constraints. Okay. Okay. And um so this is a this is a a a stimulus that exhibits this one of several. Okay. So these are rhombuses that come in different variants. And the idea is that the edges of the rhombus are each locally ambiguous and each one gives you a constraint line. So for the thin rhombus, you get these two constraint lines and they the intersection of the constraints is horizontal. For the thin case where they're moving horizontal in this case, you get these two different constraint lines, but the again the intersection of the constraints is still horizontal. Okay? But in some conditions um what the direction that you see um is actually kind of closer um to something that's pretty far off of the intersection of constraints. So I'll I'm going to show you that. Okay. So okay and the thing that we're going to manipulate here is the contrast. Okay. And so the the contrast is being manipulated here because that sort of um helps to control the degree of ambiguity of the local motion um information. And so this thing is always going to be um actually physically moving horizontally. Okay. Um and in some cases like now you probably see it as moving horizontally but we are going to manipulate some things making it thinner and thinner and thinner. Okay. Um and you will you'll actually start to see the thing kind of moving diagonally and kind of downwards. Okay. All right. So even though it's physically moving horizontally like you're seeing it as moving a little uh in a different direction moving downward. Is that working for people? Yeah. Okay. Good. All right. So that's the sort of puzzle. Yes. >> I was just going to ask if you can make it dark to see if the light effect changes at all. >> My pleasure. >> Oh wow. That's cool. >> Yeah. Yeah. So it's actually moving horizontally. Yeah. Thank you for asking. Um okay. So that's the that's the puzzle. And there were kind of as I said there were like a lot of demonstrations kind of of of that nature. And so the question is how can we understand that? All right. So to try to make sense of this um I want to go back to thinking about perception as probabistic inference. Right. So this is a slide that you have seen several times. Um so remember we think of [clears throat] the task of perception as inferring the most likely state of the world given the sensory data that you observe. Okay. So really the quantity that we're interested in is what's called the posterior. So that's the probability of this variable here which indicates something about the state of the world given an observation that's O. Okay. So the idea is that there's lots of possible states of the world. You observe O and you want to figure out the most probable state of the world given O. It's the general framework think about perception. All right. So Baze rule tells us that this posterior probability is equal to this. Okay, so there's three terms here. Okay, there's the prior, right? So that's how likely the different states of the world are in the absence of any evidence. Okay, then there's the likelihood, and that's the probability of the observed data given different hypotheses. So that's kind of how consistent the observed data are with what you would expect given different possible states of the world. Okay. And then there's a normalization factor which is the probability of the observation. Um and for a given observation that would be fixed. Okay. So then this is sort of an example of how different possible perceptual interpretations can vary in the prior and the likelihood. Right? So you have like this for this observation this particular image. The idea is that we intuitively think of this as being like a good explanation of the image, right? That there's sort of a background with these two colors and then a letter. Um, but you could also account for the image with this state of the world, right? These two objects, one that looks like that, one that looks like that. And that perfectly accounts for the data, but it's opriori very unlikely. So the prior is very low here. But the likelihood is perfectly reasonable. On the other hand, in this particular case, the prior probability of this state of the world is perfectly fine because it's like an intact letter and a reasonable looking surface, but it doesn't account for the image data. So the likelihood would be low. Okay, so a good explanation is something typically that's got high prior probability and also high likelihood. And so that's the key idea is that the the posterior the thing that we're trying to maximize um or somehow optimize when we're doing perceptual inference involves these two quantities. Okay. So that's the general framework. How can we think about this in the in the the context of motion? Yeah. All right. So um so I want to walk you through this in detail just because it's a nice kind of concrete application to these ideas that otherwise can seem kind of abstract. All right. So remember there are these two ingredients in Beijian like accounts of anything that kind of matter the likelihood and the prior and those get combined to yield the posterior. So let's talk about the likelihood. All right so remember the likelihood is the probability of the observed data given a hypothesis about the world. So we're dealing with motion perception and so now the data is going to be an image sequence. Okay and the simplest case it's just two images frame one and frame two. All right. And something moves between frame one and frame two. All right. The hypothesis space because we're dealing with motion is going to be velocity. Okay. So we're going to think about we're going to pretend that there's like only one thing moving in the world. Okay. It's defined by one velocity that is a two-dimensional vector. Okay. So it's a point in a 2D space. There's a x component of the velocity and a y component of the velocity. Okay. All right. So here's our stimulus. Okay. And we want to look at the likelihood for one particular measurement that's being made um on that uh image. And in this case, it's going to be an edge of this moving diamond. Okay? Where that circle is. All right? Okay. So the image sequence is what's ever observed, which in this case would be two frames of that moving diamond. For a given image sequence, the likelihood is a function of the unknown velocity. Right? All right. So, we're going to we're going to imagine all right for all po for every possible velocity, how probable is the observed image sequence in that through that little aperture. Okay. Okay. And so what you get here for this particular stimulus is this. Okay. What is that? Well, it's a constraint line. Okay. And what this tells us so the and I should say that the gray level here represents probability. Okay. Okay, so likelihood is the probability of that image sequence given a velocity which is a point in that 2D plane. Okay, and so what that is saying is that for all velocities that are kind of on that line, the probability is pretty high. Okay, you can also see that the line is kind of blurry. Okay, so what that means is that as you move off of the line, there's this graceful degradation in the probability of the velocity. So there's sort of a range of velocities. So if you're right on the the hot spot of the line, you know, the probability is pretty high and then you move off of that and it goes low um and then goes down to something close to zero. Okay? And so what that is saying is that there is this family of hypotheses that are all consistent with the image sequence, right? And that's because of the aperture problem and the intrinsic ambiguity of the local evidence there. Okay? So now what's manipulated here and this is actually because of this projector you this is not going to be very clear to you okay but what is what you're supposed to see as you go from this to this to this is that diamond becomes lower contrast so I think in fact maybe you can't see this at all but in fact there's a low contrast diamond here and there's an even lower contrast diamond here all right and the consequence of decreasing the contrast um is that the sensory evidence becomes noisier right so there's always noise in the measurement ment process and when the contrast is lower the image data is kind of less clear. All right. So when you kind of increase the noise or decrease the contrast what happens? Well this is what happens to the likelihood function. Okay. So we still kind of have this line here but the line is fuzzier right? So because the image data is like not as diagnostic, right? there's now like more hypotheses that can kind of account for the for the image data reasonably well. Okay. And if you decrease the contrast a lot, this things get gets really fuzzy. Okay. All right. So that's the likelihood, right, at one location. So if we're going to do Beijian inference, right? So we're going to try to to infer the the most likely velocity given this observed um image sequence. Well, we have to take our likelihood and multip by multiply it by the prior. Okay? And in particular, we also have to take the likelihood at different locations and combine those. Okay? All right. And so this is how this works, right? So we've got our image. We're making measurements at a bunch of different locations. Um again, the assumption here is that there is a single thing in the world with a single velocity. So we're just inferring this one velocity. Okay. So we've got likelihoods at different locations. This is the high contrast case where the likelihoods are these very crisp, thin, not very fuzzy constraint lines. Okay, so we've got one at this location, one at that location, and then we've got a prior. All right. And so what we're going to do is multiply these things together, and that's going to give us the posterior. All right. So the posterior is again a function of velocity. It's telling us the probability of different velocities given an image sequence. Okay. All right. So what happens? Right. Well, in this particular case when the contrast is really high, the likelihood is very peaked. Okay. And so the consequence of that is that when you multiply the likelihood at these two different locations, there's you essentially get a dot, right? You get a point, right? Because these are really thin lines. And so it's pretty much just this intersection of the constraint lines where the probability is significantly above zero. All right? So then you multiply by the prior and the prior in in this case and we'll get back to this is just assumed to be this kind of Gaussian blob here. Right? Um and you again kind of get this one point that's kin
Original Description
MIT 9.35, Spring 2024
Instructor: Josh McDermott
View the complete course: https://ocw.mit.edu/courses/9-35-perception-spring-2024
YouTube Playlist: https://www.youtube.com/playlist?list=PLUl4u3cNGP62-9RweyYBIpkqfo5dfcuS8
Continues discussion of how the brain uses input from cones to draw conclusions about the world.
License: Creative Commons BY-NC-SA
More information at https://ocw.mit.edu/terms
More courses at https://ocw.mit.edu
Support OCW at http://ow.ly/a1If50zVRl
We encourage constructive comments and discussion on OCW’s YouTube and other social media channels. Personal attacks, hate speech, trolling, and inappropriate comments are not allowed and may be removed. More details at https://ocw.mit.edu/comments.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from MIT OpenCourseWare · MIT OpenCourseWare · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
21. Post Trade Clearing, Settlement & Processing
MIT OpenCourseWare
10. Financial System Challenges & Opportunities
MIT OpenCourseWare
7. Technical Challenges
MIT OpenCourseWare
3. Blockchain Basics & Cryptography
MIT OpenCourseWare
19. Primary Markets, ICOs & Venture Capital, Part 1
MIT OpenCourseWare
1. Introduction for 15.S12 Blockchain and Money, Fall 2018
MIT OpenCourseWare
Chalk Radio, A Podcast about Inspired Teaching at MIT (Teaser)
MIT OpenCourseWare
Nuclear Gets Personal with Prof. Michael Short (S1:E1)
MIT OpenCourseWare
How Africa Has Been Made to Mean with Prof. Amah Edoh (S1:E2)
MIT OpenCourseWare
Making Deep Learning Human with Prof. Gilbert Strang (S1:E3)
MIT OpenCourseWare
Social Impact at Scale, One Project at a Time with Dr. Anjali Sastry (S1:E4)
MIT OpenCourseWare
Film is for Everyone with Prof. David Thorburn (S1:E5)
MIT OpenCourseWare
Lecture 12: Aircraft Performance
MIT OpenCourseWare
Lecture 3: Learning to Fly
MIT OpenCourseWare
Lecture 13: Interpreting Weather Data
MIT OpenCourseWare
Lecture 21: Weather Minimums and Final Tips
MIT OpenCourseWare
Hand-on, Minds On with Dr. Christopher Terman (S1:E6)
MIT OpenCourseWare
Part 4: Eigenvalues and Eigenvectors
MIT OpenCourseWare
Part 5: Singular Values and Singular Vectors
MIT OpenCourseWare
Part 3: Orthogonal Vectors
MIT OpenCourseWare
Part 2: The Big Picture of Linear Algebra
MIT OpenCourseWare
Part 1: The Column Space of a Matrix
MIT OpenCourseWare
Intro: A New Way to Start Linear Algebra
MIT OpenCourseWare
9. Chromatin Remodeling and Splicing
MIT OpenCourseWare
28. Visualizing Life - Fluorescent Proteins
MIT OpenCourseWare
20. Roth's theorem III: polynomial method and arithmetic regularity
MIT OpenCourseWare
8. Szemerédi's graph regularity lemma III: further applications
MIT OpenCourseWare
19. Roth's theorem II: Fourier analytic proof in the integers
MIT OpenCourseWare
12. Pseudorandom graphs II: second eigenvalue
MIT OpenCourseWare
1. A bridge between graph theory and additive combinatorics
MIT OpenCourseWare
Special Episode: Teaching Remotely During Covid-19 with Prof. Justin Reich
MIT OpenCourseWare
Spring 2020 Update from Dean Rajagopal
MIT OpenCourseWare
S1E7: Unpacking Misconceptions about Language & Identities with Prof. Michel DeGraff
MIT OpenCourseWare
Climate 101 Live
MIT OpenCourseWare
Welcome for Volunteers (for EarthDNA's Climate 101)
MIT OpenCourseWare
Learning to Fly with Drs. Philip Greenspun & Tina Srivastava (S1:E8)
MIT OpenCourseWare
Thinking Like an Economist with Prof. Jonathan Gruber (S1:E9)
MIT OpenCourseWare
2. Cyber Network Data Processing; AI Data Architecture
MIT OpenCourseWare
1. Artificial Intelligence and Machine Learning
MIT OpenCourseWare
2: Resistor Capacitor Circuit and Nernst Potential - Intro to Neural Computation
MIT OpenCourseWare
14: Rate Models and Perceptrons - Intro to Neural Computation
MIT OpenCourseWare
4: Hodgkin-Huxley Model Part 1 - Intro to Neural Computation
MIT OpenCourseWare
18: Recurrent Networks - Intro to Neural Computation
MIT OpenCourseWare
3: Resistor Capacitor Neuron Model - Intro to Neural Computation
MIT OpenCourseWare
15: Matrix Operations - Intro to Neural Computation
MIT OpenCourseWare
13: Spectral Analysis Part 3 - Intro to Neural Computation
MIT OpenCourseWare
16: Basis Sets - Intro to Neural Computation
MIT OpenCourseWare
20: Hopfield Networks - Intro to Neural Computation
MIT OpenCourseWare
8: Spike Trains - Intro to Neural Computation
MIT OpenCourseWare
7: Synapses - Intro to Neural Computation
MIT OpenCourseWare
19: Neural Integrators - Intro to Neural Computation
MIT OpenCourseWare
5: Hodgkin-Huxley Model Part 2 - Intro to Neural Computation
MIT OpenCourseWare
6: Dendrites - Intro to Neural Computation
MIT OpenCourseWare
17: Principal Components Analysis_ - Intro to Neural Computation
MIT OpenCourseWare
12: Spectral Analysis Part 2 - Intro to Neural Computation
MIT OpenCourseWare
11: Spectral Analysis Part 1 - Intro to Neural Computation
MIT OpenCourseWare
9: Receptive Fields - Intro to Neural Computation
MIT OpenCourseWare
10: Time Series - Intro to Neural Computation
MIT OpenCourseWare
1: Course Overview and Ionic Currents - Intro to Neural Computation
MIT OpenCourseWare
The Power of OER with Profs. Mary Rowe and Elizabeth Siler (S1:E10)
MIT OpenCourseWare
More on: RAG Basics
View skill →Related Reads
📰
📰
📰
📰
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to · Imus
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to AI
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to · Imus
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%
Dev.to AI
🎓
Tutor Explanation
DeepCamp AI