Stanford Seminar - Control-Oriented Learning for Dynamical Systems
Key Takeaways
The video discusses control-oriented learning for dynamical systems, focusing on synthesizing stabilizing feedback controllers with learned nonlinear models, and introduces techniques such as control contraction metric (CCM) and linear control design with quadratic costs. It highlights the importance of considering stabilizability and controllability in model learning and control design, and explores the use of meta-learning to learn shared structure and features for adaptive control.
Full Transcript
so um I'll get started title my talk control into learning for dynamical systems and with this sort of will revolve around is my work at the interplay between learning and control particularly for robotics so with this sort of motivated by is that if we look at each of these robotic systems on their own at least from a Dynamics perspective we're pretty good at coming up with some sort of first principles model of each of these systems um where sort of the complexity arises is where these robots are operating in different domains and under certain conditions like icy roads and contact and under different wind conditions and these are the things that are difficult to model analytically and sort of these are the reasons why we sort of turned to data-driven learning where we can't write down the model by hand so we look to data and try and figure out a model for the system overall or at least some addendum to the model some model we already have but a lot of the fundamental and sort of naive ways you might think about doing learning from machine learning don't always work very well in practice when you're talking about closed loop control so for example naive regression where you just minimize you know two Norm less may work well in minimizing one step open loop error um but control tasks happen over long time Horizons in closed loop so regression what you might find is that regression for the system and then deploying a learn model for the system with that was learned with naive regression would diverge over time with compounding errors over multiple time steps during closed loop control so what we need to do at least what I argue in this talk is that we need to orient the learning problem around what we're ultimately interested in which is closed loop control or the downstream control objective is the terminology that we use so rather than just the model fitting objective so the idea should we sort of call control oriented learning is that we should condition the model to perform better and be more adaptable in closed loop right in comparison to models that are learned naively to minimize openly error so today I'll discuss three different lines of work that we've developed time permitting to orient learning around inclusive control the first one involved translating the concept of stabilizability or controllability into algebraic conditions that we're going to use as a regularizer for closed loop control performance during offline learning the second take is a little bit different it involves structuring our model appropriately offline in a manner that's generally applicable but it naturally yields a useful control close-up controller based on lqr and then the final work I'll talk about is Shifting more towards the online setting and specifically using metal learning to learn a good extension to a nominal model that you specify maybe through first principles and learning an extension to it that we can adapt well online for adaptive control okay so I'll get started with the first project which was joint work with some folks at Google and MIT and it's going back to this regression problem and what we want to do is integrate some notion of stabilizability or the ability to stabilize the system as a constraint with the key idea being that what this should do is prune the hypothesis base of functions that were were fitting over that we're searching over and finding one that's somehow attuned to closed loop control so the key question so in stabilizability what I mean by that is there exists a controller that as a function of my current state where I am where I want to be X star and ensures that X of T my current state goes towards my X star of tea over time that's what I mean by stabilizability in the general sense okay the the key thing we need to sort of overcome here is if we want to try this out is we have to be able to describe stabilizability algebraically to construct an optimization problem that we could solve to fit a model so to do this we turn to something called contraction Theory so if we look at let's consider control alphine Dynamics x dot equals f of x plus b to x times U um and what contraction Theory starts off by doing is linearizing the Dynamics but maintaining the dependence on your linearization point so you look everywhere State and input wise and you look at the local linear Dynamics or the variational Dynamics to be more specific this is a family of linear systems and it's an exact description of small variations of your Dynamics near your linearization point and this is a linear system at least in your variations Delta X and Delta U so you can do is we can describe we can construct this notion of local energy or local distance and this is a square distance but it's weighted by some a metric called M of X it's a positive definite function Matrix function it basically generalizes the notion of distance and allows it to depend on your linearization point so like I said we're sort of generalizing notion of squared distance with some some quantity that allows it to depend on where you're linearizing it um and this thing that we're waiting the distance with is what's known as a control contraction metric or CCM and the idea is that we have a quadratic uh notion of energy we have a linear system so anybody who's done any sort of control theory that should scream out to you do some sort of linear control design that that's sort of our bread and butter as control theorists is linear control design with quadratic costs or quadratic energy functions or lyopano functions so um the idea is that if we can design a linear controller a local linear controller variational controller that's linear in the variation in x and make and design in such a way that this differential energy is decreasing locally everywhere okay well if it's decreasing locally every year and I integrate that differential control Delta U between two points x and x star so long shortest pass between x and x star I back out another control and what through some a lot more math than I intend to get into in this presentation we can show that by integrating out control for the original system we can get something that's stabilizing for the original system and it shows it makes basically makes X of T go towards X of T So to the metric M of X together with this energy decrease condition which I haven't really put any specifics on yet um together forms a certificate for stabilizability of our system so we have Dynamics and we've introduced this additional function along with some conditions that need to have satisfied to make the system stabilizable so it's a so um without going into two any detail really about how this is derived the idea that there's a few things I want to point out is that it depends on the Dynamics which go into now our model fitting objective at the top and we also have this decree condition that ends up looking like a matrix inequality that depends on both the Dynamics and the well the inverse metric okay so you've added an extra thing that you're trying to learn and it acts as a certificate for the stabilizability of the thing your the Dynamics functions that you're fitting today okay um the key problems still are those two problems one is that the constraints are infinite dimensional much like a lyapanov type analysis you want it to hold for all X um but we can't attractively enforce this as a constraint unless in uh other than in some certain conditions so what we do is we sort of relax this and say sample some X points and turn this into a finite number of constraints and there's no there's no data associated with these constraints there's no output or labels for this data it's I can sample as many x's in the state spaces I want as so I can cover the state space as densely or as sparsely as I want the only labels that get sort of show up are in the actual model fitting objective so the other intric the other problem here is that it's non-convex you have Matrix inequalities that multiples of F and W but it's by convex so what we do to solve this we linearly parametrize our metric and Dynamics and alternating and do alternating convex solves until it converge to some local Optimum um and so we actually did this on Hardware with real data uh with a quadrotor and on the left we've done sort of naive model fitting and on the right we've done our approach and what we find is that um the one on the left is not going to do so well it's going to start oscillating a bit very great here and then it starts to screw it out go to control and ends up crashing while this one on the right manages to sort of track the figure eight that it's that we've sort of commanded it to track and keep in mind this is sort of a quite a remodel that sort of learned from scratch so it's sort of it's we're not providing any base model for it sort of wanted to just see how well does this method work for learning a model from scratch okay um so with that I'll move into this or the second project which sort of takes a different spin on things and it'll still be an offline model learning and this was for some folks at MIT but it sort of takes the approach instead of trying to come up with an addendum like an extra function an extra certificate function it looks at how how do we do control or how do we build controllers for actual physical systems well often they have some structure to them that yields a controller in some way so for a mechanical system for example it has this nice LaGrange form that for example you could do a feedback linearizing control on it and it depends on the you know the second order form is specific to yielding that controller something that's a little more fundamental are linear systems when we have linear systems we have a lot of good like I said bread and butter for control theorists linear systems there's a lot of theory in designing stabilizing controllers for them specifically we can you know specify a quadratic cost and do lqr control which is optimal for linear systems so it's this idea of structure and forming control design that we're sort of latching onto here um it turns out that this linear like structure can also be identified for nonlinear systems so uh if we have f of x plus b of x times U again let's say that we can always Factor this it's fundamental theorem of calculus because Factor this as f is 0 plus a of x times x so it's some constant offset plus all the terms in your Taylor expansion that you can always factor out an X from and for such systems okay well this looks like a linear system so we can do you know we can solve a ricotta equation at each feedback timestamp for each X we encounter and do lqr like control for it it turns out practice this has worked pretty well and can do even better than lqr control for nonlinear systems um so there's a similar factorization exists for aerodynamics for any control affid system the idea of being if you write the Dynamics of X tilde dot X tilde is x minus X star so where you are minus where you want to go it turns out you can factor in a similar way where you have some Matrix function times x tilde plus b of x times U tilde and you can solve a state what's called a state dependent ricotti equation and do feedback in a non-linear manner using linear like tools and it works uh pretty well in practice there's some I think there's some analogy with what we talked about before with the certificate-based feedback with contraction Theory so um CCM based or contraction theory-based feedback you as long as you have a decreased condition satisfied you're guaranteed to be exponentially stable it's a guarantee where this bottom line here is just a decreased condition I've written out for the closed loop system with some controller pi and without getting you don't have to sort of know every term in here but the critical thing is to notice that it depends on State and input samples it depends on your Dynamics a controller and a metric so this you know three different things you're learning Dynamics controller and some certificate function it's a lot of things to try and fit to data in contrast this ricotti equation based feedback it's a heuristic it's locally optimal it's no guarantees on stabilizability but it does work well in practice and you what you're doing is you're only depending on some factories form the Dynamics and then you're doing ricotti equation based control you're not learning a separate certificate function you're not learning a separate controller the the structure the Dynamics yields your controller design so I think so what prior work has done is they've what they've tried to do is let's try and let's take this regression loss let's take this certificate decrease condition turn it into an auxiliary loss term and try and fit all these functions to data at the same time what we propose is try to extract the structure out of the Dynamics just learn the Dynamics and the structure that enables record equation based control with an auxiliary loss that sort of ensures that these factorizations are um are valid or are consistent and so what we're learning here is the Dynamics and factorizations of the Dynamics so instead learn that in it focus on learning Dynamics and some details about them and do ricotti equation based control so in a sense that you're learning less but the hope is that it still yields an effective controller so this is precisely what we did so we did in simulation we did this planar quad example we wanted to track this double Loop trajectory and going from left to right what we've done is it's an increase in the data set size so we're going to start we're going to look at how much data do we need for each different type of method to do well and on the bottom I'm going to plot the the closed loop trajectory error so this is a baseline linearization based lqr so linearizes around your current Point solver a quad equation and apply that control okay so this is the Oracle so to speak with the known Dynamics it does pretty well um then we did so this method is basically fit your Dynamics to data just do naive regression and do linearization based lqr on that model and we find it takes a while before it starts to do well we then tried out the CCM base the basically the the idea of learning Dynamics controller and certificate and again also struggles a bit to get up to the same level of performance we do our method it does really well even for a low amount of data this is purely learning the Dynamics and all at the same time learning some structural Dynamics to enable non-linear lqr like control so we found it's it's very it enables very data efficient learning and in this plots we did it through two different systems and along the x-axis is again the data set size and on the right on the upper vertical axis is sort of RMS trajectory error and we've taken sort of multiple random seeds and figured out the distribution in this these results and we get similar case where as you increase the amount of data you could do better but our method in green already does really well for low amounts of data and in fact for enough data it beats out the oracle and the key thing here is to remember that we're learning a controller that's inherently non-linear while the Oracle is using a linearization-based control so remember I said that non-linear base control we hope it does better than linearization based control and it does in a lot of cases so we can find even now we're learning the Dynamics we still get that happening so overall what we've seen is that enforcing this factorized structure enables a type of control and outperforms is more sample efficient than naive and certificate-based methods I've got a couple minutes and I'll sort of I might have to go a little fast this next section forgive me but we'll move into sort of the final section which should be more about the Adaptive than online setting so last two we're about learning a model offline and deploying it online but like I said often we're pretty good at driving at least some basic model by hand and then maybe what we want to do is um so fine tune it online so here we have a typical feedback loop from state to controller to system and then you have your feedback Arrow going backwards taking the state back to the input to the controller for an arm it could be desired your state is or position velocity or joints um and usually in a controller like I said we're relying on some knowledge of the system and the system model that we're using consists of maybe two things both structure and parameters structure could be the fact that Force equals mass times acceleration parameters are maybe the mass and shape of your payload that you maybe those are uncertain especially if you're doing pick and place operations your mass of your end effector and the inertia that you're dealing with are changing um so what adaptive control does is it says okay well maybe we don't know the value of these parameters but we maintain estimates and try and change them on online and learn them online and what's interesting is that a defining characteristic of adaptive control is that it can guarantee tracking convergence but parameter convergence is only guaranteed as a secondary result so you know parameters don't have to converge but you'll get tracking convergence a little odd the theory works so the idea is that it this is sort of naturally a a form of control oriented learning or adapting parameters on a need to know basis for control so the tricky part is maybe we don't know this structure so maybe for an arm we can write it but for a drone in wind conditions it's hard maybe to derive the interaction with the Drone and aerodynamic forces so the idea is that what we propose a do is learn this model a model for these forces and the Adaptive part is maybe just the last layer of this deep model that we want to learn and I think I'm running out of time but the idea is that if we have data from a drone from multiple different days when we're flying it or multiple different conditions that we fly it in we can treat each one as a task if we use the meta learning vernacular and the idea is that the the Dynamics the forcing dynamics that we don't know are different for each condition the idea of meta learning is learn some shared structure some shared features for each of these different conditions and adapt some parameters online that are specific to the task so this is the concept of meta learning and what prior work has done is done this with a regression objective basically differentiating through least squares or a data fit objective and we do it instead with a a control objective so we actually simulate the system and then back propagate through a simulate closed loop simulation instead of just back propagating through at least scores objective and I'll skip through but it turns out so our method or sorry the baselines are in orange and green and our method is in blue so you can see a drastic difference of performance there I apologize for rushing at the end but uh I'll turn it over to questions I guess now so if you have any questions including in the last part go for it [Applause] yes yeah very interesting talk I was just wondering um are you testing on like quadcopters throughout and is it like dependent on how linear the Dynamics of the Drone is the results you're showing here yeah so it's on a quadrupt like for this project it was just the the one system on the previous product the middle project I don't know which project do you mean specifically just in general in gen I'm just seeing always the picture of the the Drone I was just wondering how like if you had done this on the helicopter for example how well would they transfer I think that method is pretty um agnostic to the system if you look previously we did it in two different systems where so the difference between some of the baselines was not so drastic for this system because it was closer to being linear yeah well this one was a lot higher more highly non-linear and non-minimental phase so they're they're the diff the Delta between our method and baselines increases when with a nonlinearity okay very cool yep awesome talk Spencer that was really cool um I was just wondering I might have missed this but in the beginning when you're you're regressing over models F to learn is there a certain a priori structure you impose on that class of functions the first project yes yes yeah so the first one is control affine f of x plus b of x times U um there are there's also some with contraction Theory there are some assumptions about the system that you would that make it simpler to apply these um inequality conditions to construct them one of which is some sparsity in that b of X which for example from mechanical systems we hope that we can divide our actuated degrees of freedom and are unactuated degrees of freedom and it comes up with some sparsity and B of x the way awake there's one assumption involving that that we make to simplify the construction of these conditions okay so so I guess you have to have some notion of what the actual state of the system is before you begin with this yeah um so discovering a state is is not what we do okay yeah gotcha all right sounds good no I just have my own questions um I was wondering when you learn Dynamics and you constrain them to be controllable do you have any insights on what would happen if the Dynamics that you're trying to learn are not actually controllable or maybe there is some subset of the states that are controllable and some that are uncontrolled are there ways to extend your methods to those settings or what do you think would happen when you try that I think you would have to you have to find some way of dividing those things out I mean in linear systems like I'm just spitballing ideas now at this point the linear systems you can look at there's ideas like common decomposition where you can break down the controllable States from the uncontrollable States so you could look at imposing that kind of structure I think much like in model order reduction you'd have to maybe specify how many states are uncontrollable like a much like picking the lower dimensional uh the lower dimension of your model um non-linear analogs I'm not sure about but that's that approach I would go with something like that
Original Description
June 2, 2023
Spencer M. Richards of Stanford University
Robots are inherently nonlinear dynamical systems, for which synthesizing a stabilizing feedback controller with a known system model is already a difficult task. When learning a nonlinear model and controller from data, naive regression can produce a closed-loop model that is poorly conditioned for stable operation over long time horizons. In this talk, I will present our work on control-oriented learning, wherein the model learning problem is augmented to be cognizant of the desire for a stable closed-loop system. I will discuss how principles from control theory inform such augmentation to produce performant closed-loop models in a data efficient manner. This will involve ideas from contraction theory, constrained optimization, structured learning, adaptive control, and meta-learning.
Learn more about Spencer: https://stanfordasl.github.io//people/spencer-richards/
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Stanford Online · Stanford Online · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Statistical Learning: 13.2 Introduction to Multiple Testing and Family Wise Error Rate
Stanford Online
Statistical Learning: 13.1 Introduction to Hypothesis Testing II
Stanford Online
Statistical Learning: 12.R.3 Hierarchical Clustering
Stanford Online
Statistical Learning: 12.R.2 K means Clustering
Stanford Online
Statistical Learning: 12.R.1 Principal Components
Stanford Online
Statistical Learning: 13.R.1 Bonferroni and Holm II
Stanford Online
Statistical Learning: 12.6 Breast Cancer Example
Stanford Online
Statistical Learning: 12.5 Matrix Completion
Stanford Online
Statistical Learning: 12.4 Hierarchical Clustering
Stanford Online
Statistical Learning: 12.3 k means Clustering
Stanford Online
Statistical Learning: 13.1 Introduction to Hypothesis Testing
Stanford Online
Stanford Seminar - Introduction to Web3
Stanford Online
Stanford Seminar - Designing Equitable Online Experiences
Stanford Online
Stanford CS330: Deep Multi-Task & Meta Learning I 2021 I Lecture 1
Stanford Online
Stanford Seminar - Perceiving, Understanding, and Interacting through Touch
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 2
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 3
Stanford Online
Stanford CS330: Deep Multi-Task & Meta Learning I 2021 I Lecture 4
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 5
Stanford Online
Stanford Seminar - Evolution of a Web3 Company
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 6
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 7
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 8
Stanford Online
Stanford Seminar - Designing Human-Centered AI Systems for Human-AI Collaboration
Stanford Online
The Sh*tFixers: Bob Sutton Interviews David Kelley, Design Thinking Superstar
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 9
Stanford Online
Women Rise: Sheri Sheppard
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 10
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 11
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 12
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 13
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 14
Stanford Online
Stanford Webinar - Cloud Computing: What’s on the Horizon with Dr. Timothy Chou
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 15
Stanford Online
Stanford Seminar - Multi-Sensory Neural Objects: Modeling, Inference, and Applications in Robotics
Stanford Online
Stanford CS330: Deep Multi-task & Meta Learning I 2021 I Lecture 16
Stanford Online
Stanford Seminar - Toward Better Human-AI Group Decisions
Stanford Online
Stanford CS330: Deep Multi-Task & Meta Learning I 2021 I Lecture 17
Stanford Online
Stanford CS330: Deep Multi-Task & Meta Learning I 2021 I Lecture 18
Stanford Online
Stanford Webinar - Web3 Considered: Possible Futures for Decentralization and Digital Ownership
Stanford Online
Stanford Seminar - Ethics Governance-in-the-Making: Bridging Ethics Work & Governance Menlo Report
Stanford Online
Stanford Seminar - Towards Generalizable Autonomy: Duality of Discovery & Bias
Stanford Online
Stanford Seminar - ML Explainability Part 1 I Overview and Motivation for Explainability
Stanford Online
Stanford Seminar - ML Explainability Part 2 I Inherently Interpretable Models
Stanford Online
Stanford Seminar - ML Explainability Part 3 I Post hoc Explanation Methods
Stanford Online
Kratika Gupta talks about Stanford's Product Management Program
Stanford Online
Stanford Seminar - Making Teamwork an Objective Discipline - Sid Sijbrandij CEO & Chairman of GitLab
Stanford Online
Stanford Seminar - ML Explainability Part 4 I Evaluating Model Interpretations/Explanations
Stanford Online
Stanford Seminar - Adaptable Robotic Manipulation Using Tactile Sensors
Stanford Online
Stanford Seminar - ML Explainability Part 5 I Future of Model Understanding
Stanford Online
Meet Joe Lapin, Innovation and Entrepreneurship Program Completer
Stanford Online
Stanford Seminar: Social Media Scrutiny of Frontline Professionals & Implications for Accountability
Stanford Online
Stanford Seminar - Alphy and Alphy Reflect: creating a reflective mirror to advance women
Stanford Online
Stanford Webinar - The Digital Future of Health
Stanford Online
Stanford CS229M - Lecture 1: Overview, supervised learning, empirical risk minimization
Stanford Online
Stanford CS229M - Lecture 2: Asymptotic analysis, uniform convergence, Hoeffding inequality
Stanford Online
Stanford CS229M - Lecture 3: Finite hypothesis class, discretizing infinite hypothesis space
Stanford Online
Stanford Seminar - Decentralized Finance (DeFi)
Stanford Online
Stanford CS229M - Lecture 4: Advanced concentration inequalities
Stanford Online
Stanford Seminar - Bridging AI & HCI: Incorporating Human Values into the Development of AI Tech
Stanford Online
More on: ML Maths Basics
View skill →Related Reads
📰
📰
📰
📰
Decoding the Link Between Pretraining and Reinforcement Learning
Dev.to · Pneumetron
Learning the Loop — #101 | What Is Machine Learning, Really?
Medium · AI
Learning the Loop — #101 | What Is Machine Learning, Really?
Medium · Deep Learning
The model benchmark is not your production benchmark
Dev.to · hefty
🎓
Tutor Explanation
DeepCamp AI