BEYOND MAMBA AI (S6): Vector FIELDS
Key Takeaways
The video discusses going beyond the limitations of MAMBA S6 State Space Model by integrating self-attention in calculating long sequence linear compute complexity and improving in-context learning from few-shot examples with prompt engineering, using tools like Transformer, self-attention, fluid equations, and flow map.
Full Transcript
hello hello hello community so great to be back today we will go beyond mamb S6 so I will show that we can view Transformer as a flow map in a space of probability measures and we will dive in in a second just give me one minute to do a recap I did a video on member S6 and I ask hey is this an architecture that is better than the self mention we have in our Transformer networks that Empower jet GPT or Google's art and I was asked for some additional literature I like this one and I like those two YouTube videos by Albert goo and if you have to choose I would say go with the medical EI group presentation beautiful there's a beautiful remark from one of my yours and he said hey you are all wrong your mathematic is wrong the name States space model here is the same but it is all different in computer science because we are talking here about time series and in my textbook of computer science there's nothing of anything of physics and we have just here definition of a p-dimensional vector Auto regression as a state equation because in my video I used here a non-conventional approach here to the member Network I presented this as a theoretical physics and this is beautiful because this shows us exactly what's going on right now so let's dive into this for a second in this textbook there is a definition that a state space model is characterized by two principle a hidden Lattin process and we have an observation and those observation independent of the state that's it so you see this enables us to have a physical interpretation and an interpretation in computer science or time series analysis and if you look in your textbook you see that in the introduction it says hey all this here the state space model was introduced in Colman and uh buoy 1960 61 and the model arose in the space tracking setting where the state equation defines the motion equation for the position or the state of a space graft with a location X and some data y reflecting information that can be served from a tracking device such as velocity and Asim mod now if you would have gone into Google yes for my younger viewer Google is where we had this before GPT you would understand that the calman filter was developed in the 1960s for the context of the space c tracking and Aerospace particular for the Apollo project to go to the moon this common filter this state space model was an application crucial of the trajectories of man spacecraft traveling to the moon and back to Earth and we have this complicated intertwined force field after gravitational force field of Earth and the moon so you see because we do not have the possibility to solve the analytical equation in the 1960s this state space model was invented to be able to calc Cate here the trajectory of the Apollo spacecraft so you see there is a physical reference so when you say hey this is just computer science and has nothing to do with physics well you know there are some beautiful correlations here great so now update we go beyond Mamba please fasten your SE Bel seat belts because here we go you know mathematics physics dynamical system is a system that evolves over time according to a very free set of rules this is a concept used in while anything that changes over time from the motion of our celestial bodies from the Apollo spacecraft to the growth of population and it's beautiful yeah traditional deep newon networks and older will remember AR rest Nets are discrete in nature the data move to the layers in a stepwise fashion however our newal ordinary differential equations propose viewing this process as a continuous process so now imagine the data smoothly flowing through the layers of a network rather than stepping through each single layer step by step so this neural ordinary differential equations are the grand F of AI and raset here if you wonder is the grand M of AI and they offered in the old times here a powerful framework for understanding and Design deep learning models so you see this happened a long time ago now Ben for you simple neural orinary differential equation example herey you notice we can do this of course in a more complex way with a free function in the simple example here this resembles a single layer neural net neural network where the input evolves smoothly over time and of course in a more complex case allows for a sophisticated evolution of a multi-layer Network and we can model here the complex pattern and dependencies in the data so this is where we had our grandfather the newal ordinary differential equation beautiful now the concept of time in this newal ordinary differential equation is interesting so we treat now the depth of the network so the number of the layers of our Network as a Tim likee variable so instead of processing the data to discrete LEL layers n n plus 1 n plus 2 N plus whatever we Define now a continuous trajectory that the data points follow so imagine we have now somewhere a source we have a new network and then we have a flow of data flow of information to the neural network beautiful this brings us to continuous time dynamical system so in a classical systemm model to Transformations where a network as a continuous time dynamical system this is done by defining a Time demend function I just showed you in the complex case where we have state we have for example a time and we have a data parameter and this describes the range of change of the ne Network States at any time T and the function f is parameterized by Theta and this is and you're not going to believe this analogist to the weight in our new network that we know and we love beautiful however now this approach open op s up A New Perspective so with now a continuous model hey we know the equation of motion now for the Apollo spacecraft orbiting Earth and going to Moon we have the theory now we do not need an approximation by a state space model now we apply the real theoretical physics theories from dynamical system to understand and improve the neural network architecture this is the main step we're going to take so we break here whatever was the state space model and we try to insert the real analytical mathematical solution of the equation of motion for example or here fluid dynamics and you might say what corner of science are we working yes and we are working in theoretical physics and finally finally we can talk about flow maps and we have now something where there's a little bit more physical context behind all of our calculations now you know that in Floy Dynamics this Flow State equation or set of mathematical equation and I show you in a second used to describe the motion of Floyd substances such as liquids and gases this equation are fundamental in the field of theoretical physics and Engineering particular brones like aerodynamics hydrodynamics Metrology and whatever and the core concept behind all of these equation is to model how the fluid properties such as the velocity the pressure the density the temperature change in space and time in all our four-dimensional systems under various conditions isn't that beautiful and we're going to use this Insight now and we have a much better much deeper understanding what's going on now you know from mathematics and theoretical physics that we have whenever we have in a system a symmetry or a conservation of a particular parameter this gives rise to law of nature we discover Something Beautiful in this system the same happens here the foundation of the Flow State equation based on the principle here conservation of mass momentum and energy so step by step the conservation of mass in our system mass is not generated or deleted it is conserved in this system with our system boundaries so this is what we call here in physics a continuity equation and it's St the mass of a fluid remains constant as it flows through a given volume in a particular defined period of time the conservation of the momentum you notice this our beautiful and world famous navor Stokes equation of fluid dynamics they describe how the momentum of a fluid partial changes due to the forces like pressure gradients external forces and viscous forces and you notice you say hey gradient h ingredient descent hey this sounds interesting and yes you're on the right track you are there and then conservation of energy how the energy is transferred and we're going to calculate all of this beautiful yeah NV Stokes you notice we have boundary conditions we have initial conditions we have to do all of this for the mathematics but I just want to give you an idea I want to give you a feeling of what is happening yeah and this is what we calculate if you go super sonic the supersonic boom of a fet this is what we can calculate with this equation so we do not need now here a state space model that says here the airplane was here this time and now one time step later the airlane is here and suddenly a discontinuity happened when we describe the system no we have now the theoretical physical formula so describing here the system the system parameters the complete system to show why uh supersonic boom happens and with can calculate this and this is now where we go beyond and for you you have to know B this equation you have to know oiless equation and of course you know Navia Stokes equation beautiful so flow Maps is a very simple concept and you will say now I give you the definition in dynamical System Theory a flow map is a fundamental concept used to describe how a state of a system evolves over time and you say hey this is exactly like the states B model we have the state of uh I don't know Apollo and then we look how it evolves over time yes you're right and we have here an ordinary differential equation yes you're right where we say hey we look at the state for example at the position of the Apollo spaceship of a system and F is then a function dictating the systems Dynamics exactly like they did in the Apollo program and the flow map fee is now defined as a function that Maps the state of a system from one time to another this is it this is the flow map idea and you will say hey this is so familiar isn't this the state space idea no because and this is not a beauty you remember that I told you with the state space model it seems that there is because there's no self attention mechanism I could not make in context learning work so whenever you try those State space model yourself if your data you will notice there is no few short learning examples that the system exacts and the performance of in context learning and all the examples you give in the prompt are not as performant to put it mildly as in our self attention Transformer Network therefore now we say and I personally I need in context learning for my information I provide in the prom to CH4 so now we look at this and say hey we integrate self attention so we break here with the state SP model and we say no if they are not able to handle in context learning no prompt added information is accepted by the system or in a later Evolution it will be accepted because the open source Community is genius but for the moment I need self attention I need in context learning in my Transformer Alternatives so I now say hey I want to have self attention and I apply now the flow map to the self attention so we have now a dynamical system on our Transformer network with self attention so we do not go from S6 to S7 no we break here completely and we go the next evolutionary step we have now self attention in the system but we treat it differently since we have the equation of fluid we can use those equation now to calculate self attention in a much more detailed and better way okay this Viewpoint allows us to conceptualize the operation of Transformer as continuous transformation of data States rather than a series of discrete steps where we have the data propagating to the first layer to the second layer and so on you notice beautiful so you might say here in my member video I showed you here that we had the discrete time State space model here a recurrent representation and for being able to have a parallel computation on our GPU and whatsoever clusters we need here for the training here a convolutional representation and you me say hey is this not somehow the inverse not really let me me give you a little bit more details about this so make it very easy this is for our people who do not have a PhD in theoretical physics self attention is now a state transformation and you know we have an input State we have an output State and in between something happens but this now is per definition of our system a self attention mechanism because I need in context learning in my Transformers so update each token state based on its interaction with all other tokens beautiful and for this we will use some fluid equation that we know so we finally bring physics into this right oh yeah mathematics too so we have an initial State this is of course a tensor or input Matrix here this is an input sequence in Vector form you know this and then we have your self attention multi head you notice a query key and values uh structures multi-ad upates the state X to a state x- this is AK to a flow map acting on the state of a system transforming it based and this is not important we have now an interaction Dynamics defined now by the real self attention mechanism but now we Define the self attention mechanism differently and will show you in a second how and the whole thing is now a flow map interpretation so we think now of the self attention as a discrete each step in a continuous flow map and you're going to say what a brilliant idea okay okay yeah okay we go now I need just to go one step further with you and then it's done that then you understand everything so we have now a detailed mathematical model of Transformers I will show you this in a second so we're treating the Transformer now as a flow map but not in a any particular space but in the space of probability measures and we have to construct this and I will show you this how we do this so Transformers as flow maps in the space of probability measures and we have to have the interaction because if this token has here a fluid flow through the network it has to have some interaction because we have self attention so we have to kind of calculate here the attention score between each and every token in a sentence for example so what we do now we have an interacting particle system and this particle system brings us now here the mathematical describable interaction and it's a mean field to make it a little bit simpler for us beautiful this is the main idea of the new approach that goes beyond M so maybe you have to understand a little tiny bit of physics but I make it so simple you're going to laugh at the end of this video so complete description if you want to have it here on one glance Transformer flow Maps Transformer flow map is used to describe how input data evolves as it passes through the layers of a network with self attention you have an input representation you have the evolution through the layers you have the self attention mechanism but now described as a dynamical system where each token state so careful how you define a token and what becomes a token in your natural language for example is updated based on a function of the states of all the other tokens we know this from the calculation here from the classical calculation of self attention and this resembles the continuous Evolution described by the flow maps in a mathematical way then we go of course for discrete layers to a continuous flow where we can apply our Flow State equations as I told you this is our grandfather our neural ordinary differential equation where this already happened hundred of years ago in the Stone AG of EI then we implement it simple a little bit of mathematics and of course remember and I will show you this we have to go with the probabilistic interpretation because you know we have our Transformer a classical self attention way is also an auto regressive system and we have their particular probability density and we do more or less the same just in a little bit different framework if this is a little tiny bit too much for you or the first one here's a simpler form okay imagine this we have an interacting particle system and this models here the Flow State in a Transformer and we use this particle system in the context of the self attention calculation and we say hey we use simply a concept from statistical physics and dynamical systems in theoretical physics now in the classical neonian interacting particle System state each particle or entity or token has its own state which evolves not only based on its own properties but also due to the interaction with other tokens or particles particle and token are the same in our view now in a transform architecture each token it might be a word or a subword unit a token can be sort of a particle whose State this is represented of course in a classical way by the embedding Vector is influenced by all the other tokens in the sequence of course we have a semantic correlation when we talk when we use word in sentences now this is analogous to the self attention mechanism where each token representation is updated based on a relationship with every other token for example in the sequence of a sentence if you translate in English sentence to a French sentence this is what's happening so we have to model now the complex dependencies and the interaction between all different tokens and the tokens within its clusters mirroring how each part of the input sequence in a trans for influences and is influenced by the other parts now you know that in addition those interacting particle systems are inherently probabilistic which aligns well with the nature of many task and machine learning including those tackled by the Transformer so they provide here natural framework such systems can easily scale with the number of tokens making them suitable for modeling sequences of varying length and common scenario in NLP TKS and we have self attention because I need here in context learning I need to provide here for rack for example additional information in the prom so I need this for my work yeah so you have here an initial condition input sequence here a little bit more in the mathematical form here have a vector embedding then you have here the evolution to the transformal layer here the self retention mechanism here in a very simple or make it more simple you just have to Flow State inter interation beautiful now I told you know quite some formulas and I wanted we have a clear view what equations are we using first we started here with very normally ordinary differential equation you know the neural Odes were Ed for resonance then we defined what is a flow map and for you exact definition then we said hey this self attention mechanism is important for us and you know how we calculate the self attention in our Transformer with our beautiful softback beautiful and we have here the the query the key and the value uh T structures matrices expect yes and a scaling Factor yes of course and then we switch to a continuous flow dynamical system transformation very easy just a function of theta that's it good stop and we have of of course the self attention and the feet forward Network integrated and then we understand this is of course a probabilistic flow map because we are not moving here in with neutronium particles but hey we had a little bit Advanced we go with probabilistic density structures this is all that you need to know five equations simple as can be okay I will show you that it's in reality it's a little touch little bit more complex but this is just to warm up okay so now we start now now here we go you know I was asked hey why do we calculate everything here on envidia gpus we do not have little Nvidia gpus in our brain when we open the skull here now and I said hey what a beautiful image so no we have here our brain this is quite fluid if you have been studying medicine anatomy and you had your first courses and you see this the first time you say e get so the brain is quite quite s it's quite yeah semi fluid so what a beautiful introduction here to FL dyamics again just to make this clear now from a different perspective but you notice I just repeat this to be absolutely clear when we move now to fluid flows fluid equation simulate Transformers Flow State now in a nonlinear case that's a tiny bit of theoretical concept we notice it's not in your notebooks neither in theoretical physics nor in computer science nor in time spher evolation this is brand new and I will show you the source in a minute but it offers us some beautiful perspective and understanding here the nonlinear Dynamics and this is what we're looking for so fluid dynamics Behavior fluid liquids gas describe by equation you no this velocity density temperature so the evolution of token and the token rep presentation that we have chosen in our particular mathematical space now throughout the layers can be parallelized in our understanding to the flow of fluid particles so we bring a little bit of real world physics into the self attention of the Transformers and of course we have to go nonlinear because you know self attention is per definition inherently nonlinear so the mapping of this non linearity to the Floyd system is now what is so beautiful and yes you know Navia Stokes equation known for the nonlinear nature and what a coincidence that I already introduced you to them again this is the same just in a different perspective so that you wrap your head around this and you say hey it is so easy my goodness so self attention as a fluid flow you notice so we have now a complex nonlinear transformation that of course within our self attention mechanism that we said hey we want to have this in our Advanced model or mathematical representation so we integrating now the flow map with a nonlinear Floy Dynamics so we have to token Dynamics and we move over to a continuous transformation you no this now just to make it clear the third time my goodness you might say but yes so the nonlinear model interaction in Transformer you know each token index with every other token in a sentence we had the self attention self attention scores are calculated and you notice multier tension calculations and we do now the same with fluid dynamics when a fluid dynamics M this is represented as the interaction forces between particles capturing the complex nonlinear dynamics of the Flow State gorgeous so I put this in a latic file and I Ask Jud GPT 4 here to give you a short representation and here you have have exact mathematical formulation 1 2 3 4 five prob realistic interpretation you have seen this before this is just with a little touch of mathematics to be a little bit more precise because some you complain that I miss out on the mathematics yes I knew sorry try to improve myself now what we are going to need now if we do the mathematics we now enter short interrupt do you know what is the potition function and do you know why we have to normalize here in the fluid Dynamic case also our attention weights so let's let's have a look at this you know the attention scores when we calculate is in a classical transform architectures these are the calculations these are the scores now you know or maybe you have seen just in the code that the ption function Z is used to normalize those Cordes so for a given query tensor Matrix whatever set is computed here in this classical transformal self attention way so what we have we have here a DOT product between the query and the key token of yes yes yes K is some specific Dimension and the partition function now sums the exponentiated scores of all the keys serves as a denominator in the soft mix function for normalizing this great and here we have it here in our self mix function we have Z our normalized thing and then we have here our beautiful attention tensor Matrix a attention this is what we calculate here with the soft Max functionality this is what you know from the classical case and now we switch over yeah okay I repeat we have a partition function Z and the self attention mechanis so this is crucial for the model to appropriatly weigh the significance of the different parts of the input signals and then we have here the normalization of our attention weights two different things that the model attention on different tokens is proportionally distributed based on the relevance to the current token being processed this this mechanism is at the heart of the transformability who handle the sequences with a long range dependencies and to contextually understand each element in this long side sequence you notice now if we now Deep dive at a classical attention mechanism this you know how to code this but let's just make it clear what it means in the mathematical definition because we have to translate this to the Floy Dynamics equation and it's easy let me show you so just warm up self attention mechanism you notice here the Matrix a where each element a i j represent the attention weight from the token I to the Token J and then we combine here all the possible combinations and this weights determine how much each token in the sequence input sequence output sequence contributes to the representation of another token you no this and now then we're going to use an interactive particle system exactly for this description so this is the simple case so the computation of the tension weight you know this we just have done this now I just write it a little bit different so you see here the numerator is fated as scil part of the quy yes you know some of the similar terms of old tokens acting here as the normalization factor and you know the denominator is our potition function beautiful but this is what we need so the complete self attention mechanisms across all tokens can be represented as a Tor a where each row corresponds to attended weights computed for a particular token in the sequence this is the classical case now I need you to really understand and I have written a l f on this and let chat GPT explain this in simpler more beautiful English words then I can do and after I fed in my latic file to gbd4 I say hey explain why P denotes the projection after observation y onto the tangent space in my latic file and here you have now a beautiful English formulated explanation and gb4 comes back and says hey P denotes here the projection of Vector y onto the tangent space t at a point x on a manifold s typically it's a unit sphere in the context of the newal network of the Transformer architecture and this protection is essential because it ensures that operation respect the contraints on our manifold and I have of course chosen here the most simply manifold here and this is the unit sphere and you know in machine learning particularly in ual network calculation the data points or a feature representation can be constrained to lie on and only on such a manifold for reason and I will show you we have to look here especially at the normalization function so we have a we have a manifold this is the unit sphere and then we have to introduce here the concept of a tangin space a mathematical tangin space if you have a PhD mathematics you say it's so easy trivial if not I make it easy for you the tangent space at a point on a unit sphere and a manifold is the space of all possible directional velocities we operating refractors in which one can move from X while remaining on this particular manifold so in simpler terms it's a linear space that touches on the manifold at X and represents all possible direction of movement from X so you want say hey this is easy to understand beautiful so why is now the projection operator needed in this space we have to project onto the tangent space projection mathematical operation that takes a vector Y and Maps it onto the tang and space t this operation is essential when the computation like the updates in anual network and you might say hey wait has this something to do with the gradient descent right congratulation yes may lead to a point moving off the manifold but we want to have a perfect calculation so we want it on the manifold we want to have it simpler to calculate so in the context the of newal networks particular those where the feature vectors are constrained to lie on this manifold and in the easiest case is the unit sphere in a high dimensional space ensuring that the updates during the training do not move the vectors off this particular manifold and this is crucial so to have a simple operation simple mathematical performance those projection operation ensure that even after some updates some runs the vectors remain with in the geometric constraint of this particular manifold so and now why we do this we do this for this formula that you will recognize this formula when I show you this in a little tiny bit more complicated form so if Y is a vector here and the projection of Y onto our Vector space can be represented here as this projection equation and you know here YX denotes a DOT product between y and x and x is the point on our unit sphere s where the tangent space is defined and this is the projection that we're looking for and this formula effectively removes the component of our Vector Y in a high dimensional space that is orthogonal to the Tang space that is not on the tongin space but we want that everything is in onto the Tang space therefore we have to do this projection so we force this Vector back in the Tang space to get the correct direction so we ensure that the projection lies in the tong space you see as easy as that yeah Wikipedia if you say hey this was a little bit too simple as some of my viewers leave me a comment hey can you be a little bit more specific yes you go to Wikipedia I think it's beautiful over there what is a Tang based normal uh mathematical operations and then you just transform it here to our case for the new network I asked gb4 to come up with some visual interpretation of a tanget space but I have to tell you okay so we have the unit spere s S3 here and there's o yeah this should be a tangent space you see this should be the vectors that here if this is X the point x we have unit sphere s this should be here one of the vectors on a tangent sphere on this plane but yeah it's it seems that gbd4 has not been programmed for mathematical formula 3D visualization but I can can take care about this right now so we go on coming back to our topic so I say hey gb4 why did I choose with Transformers here in this particular file the unit sphere is are manold now it has to do with normalization and regularization of our mathematical composition and 24 makes it so it's so beautiful formulated I would never sort of think about this so choosing the unit spere as a manifold in the context of Transformers describing your latify is the decision likely rooted in the desire to leverage certain geometric and mathematical properties of the unit sphere some key reason why the unit sphere might be chosen as a manifold is normalization we want to maintain the norm of theor so constraining the the vectors to lie on the unit sphere effectively normalizes them to have a constant Norm constant length in our ukian picture and this can act as a form of regularization you notice preventing the magnitude of the feature vectors from going too large which might otherwise lead to numerical instabilities of our system or overfitting of the complete system so this is why we normalize it and by mapping the features here onto our particular manifold and we can chuse the unit spere in the simplest case a model can encourage their representation to be distinct and well separated beautiful this is because the maximum distance between any two points and unit spere is limited the dhere which can help it distinguish different features or different token representations yes you notice then we have a geometric interpretation of the attention mechanism itself and then we have some yeah okay T4 call it mathematical convenience I would say make it easy to Cal okay then vectors are constrained to the unit sphere in the simplest case the dot product which we calculated remember here uh in Z the dot product between any two vectors Cor respond to coine S similarity of the coine of the angle between that is a measure for the similarity is a similarity measure can be particularly useful in tension mechanism for the elements of similarity between different elements is calculated using the dot product to ensure that these calculations are more about the direction or orientation of the vector rather than evolving your different magnitudes and this is what we want we have to have here a beautiful normalized structure simplified calculation work on the unit can simplify certain mathematical operations what a coincidence for instance operation would normally require explicit normalization step that we may not need to do right now as the vect is already normalized beautiful analytical traceability yes beautiful you know this summary okay here we go so at first now we're going to explain again normalization regularization and especially in the nonlinear case where we try to treat Transformers as some flow Maps between some nonlinear States and we have then to apply this to the other uh mathematical construct so normalization you notice we just went through this dividing the vector by its Norm we have now normalized where we have for example the cian norm you have about I don't know 25 different other Norms you can use so whatever is your particular task you choose to write Norm normalization shows that all vectors have the same magnitude making the system focus on the direction orientation of the vectors rather than the magnitude this is especially useful when you use the dot product to calculate here the attention mechanism where we focusing on the angle between the vectors and in Transformers we have this in self attention beautiful regularization important a technique used to prevent overfitting by our AI system by imposing constraint or penalities on the learning process when vectors are constrainted to the unit spere they can act as a form of regularization preventing overfitting of the system you have a geometric constraint you have some mathematical implementation on a Transformer with the units versus a form of realization can help in maintaining the balance between the mod's ability to fit the training data and its generalization to new unseen data this is the classical form that you know that you love and we have a conclusion beautiful careful hey watch out the protection of unit sphere and the protection to a tangent space are different and serve different purposes so the protection onto the unit sphere is for the normalization to a unit length where the projection on a tangent space has a complete different meaning and it will have a complete different meaning and you know what is the tangent space the projection of the tangent space this is the mathematical formula please remember this formula we're going to use it in a minute an additional purpose and Transformer is used when we need to perform operation that respect the geometry of the unit sphere relevant situation we dealing with changes or update to the vector representation such as during back propagation or optimizing steps in training so now finally we come to the back propagation methodology here and here we are and yeah I've provided some additional information to chat TBD and I had some discussions so I say hey when you see updates with the vector representation such just during the backrop oration by projecting onto the Tang sphere we ensure that these updates are made in a way that it's consistent with the underlying geometry of spere do you mean that the updates of backr itself are constrained by mathematical formula or what does it mean and gb4 comes back and says yes when we talk about updates during the training in the contuct of newal network such a Transformer being constrained by this mathematical formula it prefers how the updates are computed particular in the presence of geometric constraint and you know mathematics and geometry are very densely intertwined hardly anything s in mathematics you cannot present as a geometric idea or representation and of course this also holds true for back propagation and gradient descent update so let's look at backr for a second during back prual Network you notice gradients of the loss function with respect to the model parameters a computer this gradients indicate the direction which the parameters should be adjusted to minimize the loss now we have them on our tangent space and we know exactly where they are because we project them back here to the Tang space so we have now a absolute clear Direction Where We have to move here for gradient descent the data parameters are constrained to lie in the unit spare on any manifold simply applying the standard gradient descend updates might move this parameter off so therefore adjusting the vector based on its gradient m in a vector this no longer lies on the unit sphere this is a problem for the complexity of the mathematical comp of mathematic computation and to ensure that the update parameters remain within the geometric constraint of the manifold the updates are projected back on this manifold and this in the casee of unit sphere this protecting updated vectors onto the Tang space at each point and this is why I told you the whole story because now finally we have the tools we have the power to calculate this system so my goodness yes now we start with the elure this was just the warmup to make you familiar with all of this and now comes the beauty of the physical implementation of the system if you are joining this lecture at this moment be welcome to a short summary so we have fluid State equations in the flow state of Transformers so each token in a Transformer Network can be interpreted as a particle in a fluid and I showed you why it State this means its embedding Vector representation that we have chosen for a particular parametrization for a particular tokenizer evolves as it flows through the layers of the network similar how the state of a fluid particle evolves over the time but of course we have to have some interacting forces but before we go to the self attention as the interacting forces we have to describe the Dynamics and you know he then have your St equation you know everything you know we have here the forces during the self attention phase and the self attention phase can be inter interpreted as a set of interacting forces among the tokens this are are particles now the strength the direction of these forces are determined by the low parameters of the model and the current state of the token you remember there was something with the weight t yes you're right so this interaction is where the anology to fluid dynamics becomes particular important and then as I told you we have now a continuous transformation to be able to use here the flight fluid Dynamic equation so we treat the layer index and a transform as our continuous variable but the Flow State can be modeled using now simply differential equation to describe the continuous transformation of the token State this is AK can flow State equation describe the continuous evolution of fluid particles beautiful and then we end up with a flow map representation where we have here um projection into the probability space now now we come here the real mathematics here so we say what is a transformer in this new let's call it an image of a fluid map Transformer is a flow map on the unit sphere where we have an input sequence is an initial condition which is evolved to the Dynamics and now if you want this is here our main State equation kind of here but now with physics integrated so how does here our state X at a particular time T changes over time and we have here P this is now the projection and now you know why I told you everything about here the projection to the tanget space because you're not going to believe it but P denoted the projection of Y onto the tong and space of our unit sphere and you say my goodness what a great lecture it has been up until now because now I understand exactly what this mean this is nothing else than here our projection and we have here again our potition function Z and z now is a little tiny bit more complicated but you know now what it is and we have here our key query and value tensors parameter learned from the data beta is something that you don't have to care inverse temperature so now you understand we have here the partion function we have here the projection and the change over time here of our input sequence or our initial condition our starting condition is here the projection here to our tangent space then we have some normalization and then here we simply we have here a formula all of this to understand this formula but if I would have started with this you might have decided that it is too simple or maybe a little bit too challenging for you now you remember that we are here in an interacting particle system because we need to the equation of the interacting particle system as a simplified version yes so the self attention tensor structure Matrix a i g on a particular time sequence T is now defined here in the simplified way this is what we're looking for this is now a nonlinear coupling mechanisms in the interacting particle system and this stochastic Matrix a rows are probability vectors is now the self attention Matrix of the Transformer but but understanding this now as an interacting particle system that is definitely an improvement over here the state space equation the member equation and everything else the word attention stems from the fact that here our tens captures the attention the self attention given by particle I to the particle J relatively to all particles L element in the set of s dictated by The Matrix is you know the self attention you know how to calculate this you can simplify it further beautiful yes and now we come to this beautiful publication and this publication gave me all the ideas and I had a deep dive and this I just showed you the first three formula of this this is a publication by MIT massachusett Institute of Technology Department of Mathematics and MIT Department of electrical engineering and computer science different combin interesting combination by the way MIT and CS France and yeah you can tell that this is written by some professional expert in mathematics and they have now December 22 2023 a new mathematical perspective on Transformers and this they are the Geniuses that ignited this video that I show you right now because they say now hey if you now integrate here in our transform architecture to feed forward layers so the complete Transformer Dynamics now combines all of the mechanisms above with a feet forward layer you notice from our classical Transformer and now we have here an equation for the Dynamics of the system this is how the time evolution of our system is going to happen and we have here this projection that you know of you now you know to understand you integrate here everything you know what this means and this was the whole case and the whole reason why I had this little bit more elongated introduction to this because now you can read this paper you can read now a paper by some by MIT on the latest mathematical uh development in the Transformer architecture for eii research so you see now you understand this and as I told you when we have symmetry or conservation of a particular parameter we have certain rules and here as I told you we have the continuity equation and now we look at the Contin equation exactly in this case and they told us here now the vector field driving the evolution of a single particle here clearly depends on all the other particles since we have a multi n dimensional uh particle interaction structure and one can rewrite the Dynamics here in a specific form and then you get here the evolution is governed by the Contin equation so this equation describes here our if you want probability distribution this is what we were looking for for the dynamic of the system and as I told you we also have here energy the interaction of the interaction energy Factor if you want no interaction energy here in the classical sense here and you see we can calculate now all the different parameters we need for the evolution of this system in this view as a flow state of a nonlinear dynamic system with the fluid computational forms from theoretical physics so here again you have here how the system changes over time here how here the interaction energy changes over time a little bit more complicated that you are used to when you just have a PhD in computer science but never mind there is now a beauty to this system and this beauty it's a little bit more complicated on the mathematical side I summarize here this in the words of the order please go to this this is such a beautiful publication so if you really want to see what's happening in AI research how we can have a deep dive understanding Transformers in a much better way great publication summary now the the main point here is that every particle now in our case it's a token because we are going here with natural language follows the flow of a vector field which depends on the empirical measures of all the other particles so in turn the continue equation governs the evolution of the empirical measures of the particles of the tokens whose long time behavior is of crucial interest this is what we're interested in how is the time of evolution of the system when we have here these fluid V Fields so they find out doing all the calculation the their main observation is that the particles tend to Cluster so the tokens tend to Cluster to similar token clusters under this Dynamics this is what we expect here from the classical self attention and this phenomenon is of particular Rance in learning tasks such as the next token prediction we have in the classical Transformer Network when one seeks to map a given input sequence sentence of nend tokens of n words onto a given next token Auto regressive transform architecture we predict the next word this is happening here in this case now under this mathematical formulas understanding you or applying you the physics from fluid emission the input matches encode the probability distribution of the next token and it's cluster ing indicates a small number of possible outcomes so if the time evolves in the dynamic of the system the cluster gets smaller and smaller because the next word prediction gets better and better so the set of possible next word is reduced the cluster becom smaller and smaller and they say there mathematical results indicate that the limiting distribution is actually a point Mass leaving no room for Randomness which is at ought with a iCal observation and they say this Paradox and it's a little bit technical to be honest with you maybe I do a second video if you want to see this or you read the paper so this Paradox is resolved and they found a solution that there actually is a metast stable state where we have two different time scales at work so the Transformer floor appears to possess two different time scales so in the first phase of consolidation finding the next Auto regressive token structure those tokens quickly form a few clusters the process goes on the Dynamics the calculation of our inference for example goes on so tokens are reduced cluster become smaller and smaller maybe we have 1 two three clusters around here in our space while then in a second phase and they say it's a slower phase through the process of pairwise merging of these clusters all of these possible tokens of the next possible word finally collaps to a single point and we have then the next token in our autor regressive system as you know it from the classical self mechis so they say hey if we apply this dynamics that we know from fluid physics to this Transformer and we say this is now a flow State and a nonlinear dynamic system governed by the equation of fluid dynamics of theoretical physics of theor radical physics let me mention this then we get this beautiful result that this system also converges and we get actually here the next token exactly what you expect so you see it's not that the publication attention is all you need was the only way here for our self attention to come up and develop the dynamic and finally culminate here in the next token generated here because it is a generative AI system we can use here for a deep dive if we use this we understand here Dynamics much better we can now analyze what is happening inside of the Transformers in a much I wouldn't call it an easier way but in a more physical implemented way so do we have the laws of physics really guiding the temporal evolution of our system in the probability density space okay okay this this was here a beautiful Outlook so thanks a lot here to those authors genius I love this publication please if you're interested to get an idea I have given you all the tools all the understanding that you can read this now you understand exactly what's happening December 22 so this is our Christmas present here that I would like to give you here on this particular YouTube channel for 2024 enjoy it and and it would be great to see you in the new year
Original Description
Break loose from the limitations of a MAMBA S6 State Space Model. Go beyond Mamba S6, since it doesn't integrate self-attention in calculating the long sequence linear compute complexity and may perform sub-optimal in ICL (in-context learning) from few-show examples w/ prompt engineering. (given missing benchmarks with actual LLMs and their performance.)
The central theme of the video revolves around a groundbreaking perspective of viewing Transformers in AI as analogous to flow maps in fluid dynamics. This approach suggests a dynamic and continuous flow of data through neural networks, akin to fluid movements. By adopting principles from fluid dynamics, such as modeling the evolution of a system's state variables (like in vector autoregression), the video presents a compelling fusion of physical sciences with computer science. This novel viewpoint promises to enhance the understanding of complex AI systems by applying theoretical physics concepts.
From Discrete to Continuous Models in AI.
By conceptualizing data flow through neural networks as smooth and continuous, akin to fluid motion, the speaker introduces a more sophisticated and nuanced understanding of AI models. This part of the discussion highlights the evolution from traditional neural network architectures, like ResNets, to more advanced concepts like neural ordinary differential equations, underscoring a shift towards a more fluid, dynamic model of data processing in AI.
Theoretical Physics and Mathematical Frameworks in Transformer AI Models.
The latter part of the video emphasizes the application of theoretical physics and advanced mathematical frameworks in understanding and improving Transformer architectures. The speaker explores how integrating concepts such as fluid state equations and flow maps can provide a deeper, more comprehensive understanding of the dynamics within Transformer models. This perspective not only enhances the interpretation of AI behavior but also opens new avenues for opt
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Discover AI · Discover AI · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Discover AI
Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Discover AI
Create a Smarter Future!
Discover AI
The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
Discover AI
Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Discover AI
Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Discover AI
Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D (SBERT 48)
Discover AI
Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey! (SBERT 49)
Discover AI
SBERT Extreme 3D: Train a BERT Tokenizer on your (scientific) Domain Knowledge (SBERT 50)
Discover AI
Discover Vision Transformer (ViT) Tech in 2023
Discover AI
Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Discover AI
Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Discover AI
BERT and GPT in Language Models like ChatGPT or BLOOM | EASY Tutorial on Large Language Models LLM
Discover AI
Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source) #shorts
Discover AI
From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
Discover AI
How to start with ChatGPT? | Short Introduction to OpenAI API #shorts
Discover AI
The Future of Conversational AI? Google's PaLM w/ RLHF | LLM ChatGPT Competitor
Discover AI
Microsoft and ChatGPU
Discover AI
From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
Discover AI
Google's 2nd Answer to "BING ChatGPT": Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Discover AI
TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
Discover AI
3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
Discover AI
FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
Discover AI
ChatGPT - Can it Lie to you?
Discover AI
ChatGPT Alternative: Perplexity by Perplexity.AI
Discover AI
2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
Discover AI
Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Discover AI
BLOOM 176B Inference on AWS | Bigger than GPT-3 for more Power!
Discover AI
Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings? My own ChatGPT? | Visual Q+A
Discover AI
Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Discover AI
After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
Discover AI
Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Discover AI
Fine-tune ChatGPT w/ in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Discover AI
The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
Discover AI
New TECH: Vision Transformer 2023 on Image Classification | AI
Discover AI
PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned | AI Tech
Discover AI
New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
Discover AI
New BING ChatGPT loses its mind
Discover AI
Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Discover AI
Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Discover AI
Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Discover AI
PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
Discover AI
New BING Chat AGGRESSIVE
Discover AI
Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Discover AI
Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Discover AI
Dream Job Alert: AI Prompt Engineer - $335K | AI Prompt Design: A Crash Course
Discover AI
Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Discover AI
Microsoft's CEO in Trouble #shorts
Discover AI
Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Discover AI
OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
Discover AI
ChatGPT polarizes
Discover AI
Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Discover AI
ChatGPT Prompt Engineering w/ in-context learning (ICL) - 7 Examples | Tutorial
Discover AI
Chat with your Image! BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Discover AI
ChatGPT: Multidimensional Prompts
Discover AI
ChatGPT: In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
Discover AI
Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Discover AI
Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Discover AI
Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Discover AI
Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Discover AI
More on: LLM Foundations
View skill →Related Reads
🎓
Tutor Explanation
DeepCamp AI