arxiv: v1 [cs.lg] 23 May 2018

Size: px

Start display at page:

Download "arxiv: v1 [cs.lg] 23 May 2018"

Ashley Harrell
5 years ago
Views:

1 Regret Bounds for Robust Adaptive Control of the Linear Quadratic Regulator Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu University of California, Berkeley arxiv: v1 cslg 23 May 2018 May 25, 2018 Abstract We consider adaptive control of the Linear Quadratic Regulator LQR, where an unknown linear system is controlled subject to quadratic costs Leveraging recent developments in the estimation of linear systems and in robust controller synthesis, we present the first provably polynomial time algorithm that provides high probability guarantees of sub-linear regret on this problem We further study the interplay between regret minimization and parameter estimation by proving a lower bound on the expected regret in terms of the exploration schedule used by any algorithm Finally, we conduct a numerical study comparing our robust adaptive algorithm to other methods from the adaptive LQR literature, and demonstrate the flexibility of our proposed method by extending it to a demand forecasting problem subject to state constraints 1 Introduction The problem of adaptively controlling an unknown dynamical system has a rich history, with classical asymptotic results of convergence and stability dating back decades 15, 16 Of late, there has been a renewed interest in the study of a particular instance of such problems, namely the adaptive Linear Quadratic Regulator LQR, with an emphasis on non-asymptotic guarantees of stability and performance Initiated by Abbasi-Yadkori and Szepesvári 2, there have since been several works analyzing the regret suffered by various adaptive algorithms on LQR here the regret incurred by an algorithm is thought of as a measure of deviations in performance from optimality over time These results can be broadly divided into two categories: those providing high-probability guarantees for a single execution of the algorithm 2, 5, 11, 14, and those providing bounds on the expected Bayesian regret incurred over a family of possible systems 3, 21 As we discuss in more detail, these methods all suffer from one or several of the following limitations: restrictive and unverifiable assumptions, limited applicability, and computationally intractable subroutines In this paper, we provide, to the best of our knowledge, the first polynomial-time algorithm for the adaptive LQR problem that provides high probability guarantees of sub-linear regret, and that does not require unverifiable or unrealistic assumptions Related Work There is a rich body of work on the estimation of linear systems as well as on the robust and adaptive control of unknown systems We target our discussion to works on non-asymptotic guarantees for the LQR control of an unknown system, broadly divided into three categories 1

2 Offline estimation and control synthesis: In a non-adaptive setting, ie, when system identification can be done offline prior to controller synthesis and implementation, the first work to provide endto-end guarantees for the LQR optimal control problem is that of Fiechter 13, who shows that the discounted LQR problem is PAC-learnable Dean et al Dean et al 8 improve on this result, and provide the first end-to-end sample complexity guarantees for the infinite horizon average cost LQR problem Optimism in the Face of Uncertainty OFU: Abbasi-Yadkori and Szepesvári 2, Faradonbeh et al 11, and Ibrahimi et al 14 employ the Optimism in the Face of Uncertainty OFU principle 7, which optimistically selects model parameters from a confidence set by choosing those that lead to the best closed-loop infinite horizon control performance, and then plays the corresponding optimal controller, repeating this process online as the confidence set shrinks While OFU in the LQR setting has been shown to achieve optimal regret Õ T, its implementation requires solving a non-convex optimization problem to precision ÕT /2, for which no provably efficient implementation exists Thompson Sampling TS: To circumvent the computational roadblock of OFU, recent works replace the intractable OFU subroutine with a random draw from the model uncertainty set, resulting in Thompson Sampling TS based policies 3, 5, 21 Abeille and Lazaric 5 show that such a method achieves ÕT 2/3 regret with high-probability for scalar systems However, their proof does not extend to the non-scalar setting Abbasi-Yadkori and Szepesvári 3 and Ouyang et al 21 consider expected regret in a Bayesian setting, and provide TS methods which achieve Õ T regret Although not directly comparable to our result, we remark on the computational challenges of these algorithms Whereas the proof of Abbasi-Yadkori and Szepesvári 3 was shown to be incorrect 20, Ouyang et al 21 make the restrictive assumption that there exists a known initial compact set Θ describing the uncertainty in the system parameters, such that for any system θ 1 Θ, the optimal controller Kθ 1 is stabilizing when applied to any other system θ 2 Θ No means of constructing such a set are provided, and there is no known tractable algorithm to verify if a given set satisfies this property Also, it is implicitly assumed that projecting onto this set can be done efficiently Contributions To develop the first polynomial-time algorithm that provides high probability guarantees of sub-linear regret, we leverage recent results from the estimation of linear systems 23, robust controller synthesis 18, 25, and coarse-id control 8 We show that our robust adaptive control algorithm: i guarantees stability and near-optimal performance at all times; ii achieves a regret up to time T bounded by ÕT 2/3 ; and iii is based on finite-dimensional semidefinite programs of size logarithmic in T Furthermore, our method estimates the system parameters at ÕT /3 rate in operator norm Although system parameter identification is not necessary for optimal control performance, an accurate system model is often desirable in practice Motivated by this, we study the interplay between regret minimization and parameter estimation, and identify fundamental limits connecting the two We show that the expected regret of our algorithm is lower bounded by ΩT 2/3, proving that our analysis is sharp up to logarithmic factors Moreover, our lower bound suggests that the estimation rate achievable by any algorithm with OT α regret is ΩT α/2 Finally, we conduct a numerical study of the adaptive LQR problem, in which we implement our algorithm, and compare its performance to heuristic implementations of OFU and TS based methods We show on several examples that the regret incurred by our algorithm is comparable to that of the OFU and TS based methods Furthermore, the infinite horizon cost achieved by our algorithm at any given time on the true system is consistently lower than that attained by OFU 2

3 and TS based algorithms Finally, we use a demand forecasting example to show how our algorithm naturally generalizes to incorporate environmental uncertainty and safety constraints 2 Problem Statement and Preliminaries In this work we consider adaptive control of the following discrete-time linear system iid x k+1 = A x k + B u k + w k, w k N 0, σwi 2, 21 where x k R n is the state, u k R p is the control input, and w k R n is the process noise We assume that the state variables are observed exactly and, for simplicity, that x 0 = 0 We consider the Linear Quadratic Regulator optimal control problem, given by cost matrices Q 0 and R 0, J = min u lim T T 1 T E x k Qx k + u k Ru k k=1 st dynamics 21, where the minimum is taken over measurable functions u = {u k } k 1, with each u k adapted to the history x k, x k,, x 1, and possibe additional randomness independent of future states Given knowledge of A, B, the optimal policy is a static state-feedback law u k = K x k, where K is derived from the solution to a discrete algebraic Riccati equation We are interested in algorithms which operate without knowledge of the true system transition matrices A, B We measure the performance of such algorithms via their regret, defined as RegretT := x k Qx k + u k Ru k J k=1 The regret of any algorithm is lower-bounded by Ω T, a bound matched by OFU up to logarithmic factors 11 However, after each epoch, OFU requires optimizing a non-convex objective to OT /2 precision Instead, our method uses a subroutine based on convex optimization and robust control 21 Preliminaries: System Level Synthesis We briefly describe the necessary background on robust control and System Level Synthesis 25 SLS These tools were recently used by Dean et al 8 to provide non-asymptotic bounds for LQR in the offline estimate-and-then-control setting In Appendix A, we expand on these preliminaries Consider the dynamics 21, and fix a static state-feedback control policy K, ie, let u k = Kx k Then, the closed loop map from the disturbance process {w 0, w 1, } to the state x k and control input u k at time k is given by x k = k t=1 A + B K k t w t, u k = k t=1 KA + B K k t w t 22 Letting Φ x k := A + B K k and Φ u k := KA + B K k, we can rewrite Eq 22 as xk u k = k t=1 Φx k t + 1 w Φ u k t + 1 t, 23 3

4 where {Φ x k, Φ u k} are called the closed loop system response elements induced by the controller K The SLS framework shows that for any elements {Φ x k, Φ u k} constrained to obey Φ x k + 1 = A Φ x k + B Φ u k, Φ x 1 = I, k 1, 24 there exists some controller that achieves the desired system responses 23 Theorem A1 formalizes this observation: the SLS framework thereore allows for any optimal control problem over linear systems to be cast as an optimization problem over elements {Φ x k, Φ u k}, constrained to satisfy the affine equations 24 Comparing equations 22 and 23, we see that the former is non-convex in the controller K, whereas the latter is affine in the elements {Φ x k, Φ u k}, enabling solutions to previously difficult optimal control problems As we work with infinite horizon problems, it is notationally more convenient to work with transfer function representations of the above objects, which can be obtained by taking a z-transform of their time-domain representations The frequency domain variable z can be informally thought of as the time-shift operator, ie, z{x k, x k+1, } = {x k+1, x k+2, }, allowing for a compact representation of LTI dynamics We use boldface letters to denote such transfer functions, eg, Φ x z = k=1 Φ xkz k Then, the constraints 24 can be rewritten as Φ zi A B x = I, Φ u and the corresponding not necessarily static control law u = Kx is given by K = Φ u Φ x Although other approaches to optimal controller design exists, we argue now that the SLS parameterization has some appealing properties when applied to the control of uncertain systems In particular, suppose that rather than having access to the true system transition matrices A, B, we instead only have access to estimates Â, B The SLS framework allows us to characterize the system responses achieved by a controller, computed using only the estimates Â, B, on the true system A, B Specifically, if we denote := Â A Φ x + B B Φ u, simple algebra shows that zi Â B Φ x = I if and only if Φ u Φ zi A B x = I + Φ u Theorem A2 shows that if I + exists, then the controller K = Φ u Φ x, computed using only the estimates Â, B, achieves the following response on the true system A, B : x Φx = I + u w Φ u Further, if K stabilizes the system Â, B, and I + is stable simple sufficient conditions can be derived to ensure this, see 8, then K is also stabilizing for the true system It is this transparency between system uncertainty and controller performance that we exploit in our algorithm We end this discussion with the definition of a function space that we use extensively throughout: { } SC, ρ := M = Mkz k Mk Cρ k, k = 0, 1, 2, k=0 4

5 The space SC, ρ consists of stable transfer functions that satisfy a certain decay rate in the spectral norm of their impulse response elements We denote the restriction of SC, ρ to the space of F -length finite impulse response FIR filters by S F C, ρ, ie, M S F C, ρ if M SC, ρ, and Mk = 0 for all k > F Further note that we write M 1 z SC, ρ to mean that zm SC, ρ, ie, that M0 = 0 We equip SC, ρ with the H and H 2 norms, which are infinite horizon analogs of the spectral and Frobenius norms of a matrix, respectively: M H = sup w 2 =1 Mw 2 and M H2 = k=0 Mk 2 F The H and H 2 norm have distinct interpretations The H norm of a system M is equal to its l 2 l 2 operator norm, and can be used to measure the robustness of a system to unmodelled dynamics 26 The H 2 norm has a direct interpretation as the energy transferred to the system by a white noise process, and is hence closely related to the LQR optimal control problem Unsurprisingly, the H 2 norm appears in the objective function of our optimization problem, whereas the H norm appears in the constraints to ensure robust stability and performance 3 Algorithm and Guarantees Our proposed robust adaptive control algorithm for LQR is shown in Algorithm 1 We note that while Line 8 of Algorithm 1 is written as an infinite-dimensional optimization problem, because of the FIR nature of the decision variables, it can be equivalently written as a finite-dimensional semidefinite program We describe this transformation in Section G3 of the Appendix Some remarks on practice are in order First, in Line 6, only the trajectory data collected during the i-th epoch is used for the least squares estimate Second, the epoch lengths we use grow exponentially in the epoch index These settings are chosen primarily to simplify the analysis; in practice all the data collected should be used, and it may be preferable to use a slower growing epoch schedule such as T i = C T i + 1 Finally, for storage considerations, instead of performing a batch least squares update of the model, a recursive least squares RLS estimator rule can be used to update the parameters in an online manner 31 Regret Upper Bounds Our guarantees for Algorithm 1 are stated in terms of certain system specific constants, which we define here We let K denote the static feedback solution to the LQR problem for A, B, Q, R Next, we define C, ρ such that the closed loop system A + B K belongs to SC, ρ Our main assumption is stated as follows Assumption 31 We are given a controller K 0 that stabilizes the true system A, B Furthermore, letting Φ x, Φ u denote the response of K 0 on A, B, we assume that Φ x SC x, ρ and Φ u SC u, ρ, where the constants C x, C u, ρ are defined in Algorithm 1 The requirement of an initial stabilizing controller K 0 is not restrictive; Dean et al 8 provide an offline strategy for finding such a controller Furthermore, in practice Algorithm 1 can be initialized with no controller, with random inputs applied instead to the system in the first epoch to estimate A, B within an initial confidence set for which the synthesis problem becomes feasible Our first guarantee is on the rate of estimation of A, B as the algorithm progresses through time This result builds on recent progress 23 for estimation along trajectories of a linear dynamical system 5

6 Algorithm 1 Robust Adaptive Control Algorithm Require: Stabilizing controller K 0, failure probability δ 0, 1, and constants C, ρ, K 1: Set C x O1C 1 ρ, C 3 u K C x, and ρ ρ 2: Set C T n Õ + p C4 1+ K 4 1 ρ 8 3: for i = 0, 1, 2, do 4: Set T i C T 2 i and ση,i 2 σ2 wt i /C T /3 5: D i = {x i k, ui k }T i k=1 evolve system forward T i steps using feedback u = K i x + η i, where each entry of η i is drawn iid from N 0, ση,i 2 I p 6: Âi, B i arg min Ti A,B k=1 1 2 xi k+1 Axi k Bui k 2 2 7: Set ε i Õ σw K C n+p σ η,i 1 ρ 3 T i and F i Õ1i+1 1 ρ 8: Set K i+1 = Φ u Φ x, where Φ x, Φ u are the solution to 1 minimize γ 0,1 min Q 1/2 0 Φx 1 γ Φ x,φ u,v 0 R 1/2 Φ u H 2 st zi Âi B Φ x i = I + 1 Φ u z F V, 2εi Φx i 1 C x ρ F i+1 γ, Φ u H V C x ρ F i+1, Φ x 1 z S F i C x, ρ, Φ u 1 z S F i C u, ρ 9: end for For what follows, the notation T, Õ hides absolute constants and polylog 1 δ, C 1, 1 ρ, n, p, B, K factors Theorem 32 Fix a δ 0, 1 and suppose that Assumption 31 holds With probability at least 1 δ the following statement holds Suppose that T is at an epoch boundary Let ÂT, BT denote the current estimate of A, B computed by Algorithm 1 at the end of time T Then, this estimate satisfies the guarantee max{ ÂT A, BT B } Õ C K n + p 1 ρ 3 Theorem 32 shows that Algorithm 1 achieves a consistent estimate of the true dynamics A, B, and learns at a rate of ÕT /3 We note that consistency of parameter estimates is not a guarantee provided by OFU or TS based approaches Next, we state an upper bound on the regret incurred by Algorithm 1 Theorem 33 Fix a δ 0, 1 and suppose that Assumption 31 holds With probability at least 1 δ the following statement holds For all T 0 we have that Algorithm 1 satisfies RegretT Õ n + p C4 1 + K B 2 J 1 ρ 16 T 2/3 Here, the notation Õ also hides ot 2/3 terms T 1/3 6

7 The intuition behind our proof is transparent We use SLS to show that the cost during epoch i is bounded by T i 1 + Oση,i 2 /σ2 w1 + Oε i J, where the 1 + Oε i factor is the performance degredation incurred by model uncertainty, and the 1 + Oση,i 2 /σ2 w factor is the additional cost incurred from injecting exploration noise Hence, the regret incurred during this epoch is OT i ση,i 2 /σ2 w + ε i J, We then bound our estimation error by ε i = Õσ w/σ η,i T /2 i Setting ση,i 2 = σ2 wti α 1 α, we have the per epoch bound ÕTi + T 1 α/2 i Choosing α = 1/3 to balance these competing powers of T i and summing over logarithmic number of epochs, we obtain a final regret of ÕT 2/3 The main difficulty in the proof is ensuring that the transient behavior of the resulting controllers is uniformly bounded when applied to the true system Prior works sidestep this issue by assuming that the true dynamics lie within a known compact set for which the Heine-Borel theorem asserts the existence of finite constants that capture this behavior We go a step further and work through the perturbation analysis which allows us to give a regret bound that depends only on simple quantities of the true system A, B The full proof is given in the appendix Finally, we remark that the dependence on 1/1 ρ in our results is an artifact of our perturbation analysis, and we leave sharpening this dependence to future work 32 Regret Lower Bounds and Parameter Estimation Rates We saw that Algorithm 1 achieves ÕT 2/3 regret with high probability Now we provide a matching algorithmic lower bound on the expected regret, showing that the analysis presented in Section 31 is sharp as a function of T Moreover, our lower bound characterizes how much regret must be accrued in order to achieve a specified estimation rate for the system parameters A, B Theorem 34 Let the initial state x 0 be distributed according to the steady state distribution N 0, P of the optimal closed loop system, and let {u t } t 0 be any sequence of inputs as in Section 2 Furthermore, let f : R R be any function such that with probability 1 δ we have λ min k=0 xk u k x k Then, there exist positive values T 0 and C 0 such that for all T T 0 we have k=0 u k ft 31 E x k Qx k + u k Ru k J δλ minr1 + σ min K 2 ft T 0 C 0, where T 0 and C 0 are functions of A, B, Q, R, σ 2 w, and n, detailed in Appendix E The proof of the estimation error Theorem 32 shows that Algorithm 1 satisfies Eq 31 with ft = ÕT σ2 η,θlog 2 T Since the exploration variance σ2 η,i used by Algorithm 1 during the i-th epoch is given by ση,i 2 = Oσ2 wt i/3, we obtain the following corollary which demonstrates the sharpness of our regret analysis with respect to the scaling of T Corollary 35 For T > C 1 n, δ, σ 2 w, A, B, Q, R the expected regret of Algorithm 1 satisfies k=1 E x k Qx k + u k Ru k J Ωλ min R1 + σ min K 2 T 2/3 7

8 A natural question to ask is how much regret does any algorithm accrue in order to achieve estimation error Â A ε and B B ε From Theorem 32 we know that Algorithm 1 estimates A, B at rate ÕT /3 Therefore, in order to achieve ε estimation error, T must be Ωε 3 Hence, Theorem 33 implies that the regret of Algorithm 1 to achieve ε estimation error is Õε 2 Interestingly, let us consider any other Algorithm achieving OT α regret for some 0 < α < 1 Then, Theorem 34 suggests that the best rate achievable by such an algorithm is OT α/2, since the minimum eigenvalue condition Eq 31 governs the signal-to-noise ratio In the case of linearregression with independent data it is known that the minimax estimation rate is lower bounded by square root of the inverse of the minimum eigenvalue 31 We conjecture that the same results holds in our case Therefore, to achieve ε estimation error, any Algorithm would likely require Ωε 2 regret, showing that Algorithm 1 is optimal up to logarithmic factors in this sense Finally, we note that while Algorithm 1 estimates A, B at a rate ÕT /3, Theorem 34 suggests that any algorithm achieving the O T regret would estimate A, B at a rate ΩT /4 4 Experiments Regret Comparison We illustrate the performance of several adaptive schemes empirically We compare the proposed robust adaptive method with non-bayesian Thompson sampling TS as in Abeille and Lazaric 5 and a heuristic projected gradient descent PGD implementation of OFU As a simple baseline, we use the nominal control method, which synthesizes the optimal infinite-horizon LQR controller for the estimated system and injects noise with the same schedule as the robust approach Implementation details and computational considerations for all adaptive methods are in Appendix G The comparison experiments are carried out on the following LQR problem: A = , B = I, Q = 10I, R = I, σ w = This system corresponds to a marginally unstable Laplacian system where adjacent nodes are weakly connected; these dynamics were also studied by 4, 8, 24 The cost is such that input size is penalized relatively less than state This problem setting is amenable to robust methods due to both the cost ratio and the marginal instability, which are factors that may hurt optimistic methods In Appendix H1, we show similar results for an unstable system with large transients To standardize the initialization of the various adaptive methods, we use a rollout of length T 0 = 100 where the input is a stabilizing controller plus Gaussian noise with fixed variance σ u = 1 This trajectory is not counted towards the regret, but the recorded states and inputs are used to initialize parameter estimates In each experiment, the system starts from x 0 = 0 to reduce variance over runs For all methods, the actual errors Ât A and B t B are used rather than bounds or bootstrapped estimates The effect of this choice on regret is small, as examined empirically in Appendix H2 The performance of the various adaptive methods is compared in Figure 1 The median and 90th percentile regret over 500 instances is displayed in Figure 1a, which gives an idea of both typical and worst-case behavior The regret of the optimal LQR controller for the true system is displayed as a baseline Overall, the methods have very similar performance One benefit of robustness is the 8

9 a Regret b Infinite Horizon LQR Cost Regret OFU TS Robust Nominal Optimal Iteration Cost Suboptimality OFU TS Robust Nominal Iteration Figure 1: A comparison of different adaptive methods on 500 experiments of the marginally unstable Laplacian example in 41 In a, the median and 90th percentile regret is plotted over time In b, the median and 90th percentile infinite-horizon LQR cost of the epoch s controller a Demand Forecasting b Constraint Satisfaction stochasticity LTI filter demand temperature control state norm no constraint constrained Iteration Figure 2: The addition of constraints in the robust synthesis problem can guarantee the safe execution of adaptive systems We consider an example inspired by demand forecasting, as illustrated in a, where the left hand side of the diagram represents unknown dynamics The median and maximum values of x t over 500 trials are plotted for both the unconstrained and constrained synthesis problems in b guaranteed stability and bounded infinite-horizon cost at every point during operation In Figure 1b, this infinite-horizon LQR cost is plotted for the controllers played during each epoch This value measures the cost of using each epoch s controller indefinitely, rather than continuing to update its parameters The robust adaptive method performs relatively better than other adaptive algorithms, indicating that it is more amenable to early stopping, ie, to turning off the adaptive component of the algorithm and playing the current controller indefinitely Extension to Uncertain Environment with State Constraints The proposed robust adaptive method naturally generalizes beyond the standard LQR problem We consider a disturbance forecasting example which incorporates environmental uncertainty and safety constraints Consider a system with known dynamics driven by stochastic disturbances that are now correlated in time We model the disturbance process as the output of an unknown autonomous LTI system, as illustrated in Figure 2a This setting can be interpreted as a demand forecasting problem, where, for example, the system is a server farm and the disturbances represent changes in the amount of incoming jobs 9

10 If the dynamics of the correlated disturbance process are known, this knowledge can be used for more cost-effective temperature control We let the system A, B with known dynamics be described by the graph Laplacian dynamics as in Eq 41 The disturbance dynamics are unknown and are governed by a stable system transition matrix A d, resulting in the following dynamics for the full system: xt+1 d t+1 A I zt = + 0 A d d t B u 0 t + 0 w I t, A d = The costs are set to model expensive inputs, with Q = I and R = I The controller synthesis problem in Line 8 of Algorithm 1 is modified to reflect the problem structure, and crucially, we add a constraint on the system response Φ x Further details of the formulation are explained in Appendix H3 Figure 2b illustrates the effect While the unconstrained synthesis results in trajectories with large state values, the constrained synthesis results in much more moderate behavior 5 Conclusions and Future Work We presented a polynomial-time algorithm for the adaptive LQR problem that provides high probability guarantees of sub-linear regret In contrast to other approaches to this problem, our robust adaptive method guarantees stability, robust performance, and parameter estimation We also explored the interplay between regret minimization and parameter estimation, identifying fundamental limits connecting the two Several questions remain to be answered It is an open question whether a polynomial-time algorithm can achieve a regret of Õ T In our implementation of OFU, we observed that PGD performed quite effectively Interesting future work is to see if the techniques of Fazel et al 12 for policy gradient optimization on LQR can be applied to prove convergence of PGD on the OFU subroutine, which would provide an optimal polynomial-time algorithm Moreover, we observed that OFU and TS methods in practice gave estimates of system parameters that were comparable with our method which explicitly adds excitation noise It seems that the switching of control policies at epoch boundaries provides more excitation for system identification than is currently understood by the theory Furthermore, practical issues that remain to be addressed include satisfying safety constraints and dealing with nonlinear dynamics; in both settings, finite-sample parameter estimation/system identification and adaptive control remain an open problem Acknowledgments SD is supported by an NSF Graduate Research Fellowship As part of the RISE lab, HM is generally supported in part by NSF CISE Expeditions Award CCF , DHS Award HSHQDC , and gifts from Alibaba, Amazon Web Services, Ant Financial, CapitalOne, Ericsson, GE, Google, Huawei, Intel, IBM, Microsoft, Scotiabank, Splunk and VMware BR is generously supported in part by NSF award CCF , ONR awards N , N , and N , the DARPA Fundamental Limits of Learning Fun LoL and Lagrange Programs, and an Amazon AWS AI Research Award 10

11 References 1 Yasin Abbasi-Yadkori Online Learning for Linearly Parametrized Control Problems PhD thesis, University of Alberta, Yasin Abbasi-Yadkori and Csaba Szepesvári Regret Bounds for the Adaptive Control of Linear Quadratic Systems In Conference on Learning Theory, Yasin Abbasi-Yadkori and Csaba Szepesvári Bayesian Optimal Control of Smoothly Parameterized Systems: The Lazy Posterior Sampling Algorithm In Conference on Uncertainty in Artificial Intelligence, Yasin Abbasi-Yadkori, Nevena Lazic, and Csaba Szepesvári Regret Bounds for Model-Free Linear Quadratic Control arxiv: , Marc Abeille and Alessandro Lazaric Thompson Sampling for Linear-Quadratic Control Problems In AISTATS, James Anderson and Nikolai Matni Structured State Space Realizations for SLS Distributed Controllers In Allerton, S Bittanti and M C Campi Adaptive control of linear time invariant systems: the bet on the best principle Communications in Information and Systems, 64, Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu On the Sample Complexity of the Linear Quadratic Regulator arxiv: , Steven Diamond and Stephen Boyd CVXPY: A Python-embedded modeling language for convex optimization Journal of Machine Learning Research, 1783, Bogdan Dumitrescu Positive trigonometric polynomials and signal processing applications Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis Finite Time Analysis of Optimal Adaptive Policies for Linear-Quadratic Systems arxiv: , Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi Global Convergence of Policy Gradient Methods for Linearized Control Problems arxiv: , Claude-Nicolas Fiechter PAC Adaptive Control of Linear Systems In Conference on Learning Theory, Morteza Ibrahimi, Adel Javanmard, and Benjamin Van Roy Efficient Reinforcement Learning for High Dimensional Linear Quadratic Systems In Neural Information Processing Systems, Petros A Ioannou and Jing Sun Robust adaptive control, volume 1 PTR Prentice-Hall Upper Saddle River, NJ, Miroslav Krstic, Ioannis Kanellakopoulos, and Peter V Kokotovic Nonlinear and adaptive control design Wiley,

12 17 Bo Lincoln and Anders Rantzer Relaxing dynamic programming IEEE Transactions on Automatic Control, 518: , Nikolai Matni, Yuh-Shyang Wang, and James Anderson Scalable system level synthesis for virtually localizable systems In IEEE Conference on Decision and Control, Brendan O Donoghue, Eric Chu, Neal Parikh, and Stephen Boyd Conic Optimization via Operator Splitting and Homogeneous Self-Dual Embedding Journal of Optimization Theory and Applications, 1693, Ian Osband and Benjamin Van Roy Posterior Sampling for Reinforcement Learning Without Episodes arxiv: , Yi Ouyang, Mukul Gagrani, and Rahul Jain Learning-based Control of Unknown Linear Systems with Thompson Sampling arxiv: , Mark Rudelson and Roman Vershynin Hanson-Wright inequality and sub-gaussian concentration Electronic Communications in Probability, 1882, Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht Learning Without Mixing: Towards A Sharp Analysis of Linear System Identification arxiv: , Stephen Tu and Benjamin Recht Least-Squares Temporal Difference Learning for the Linear Quadratic Regulator arxiv: , Yuh-Shyang Wang, Nikolai Matni, and John C Doyle A System Level Approach to Controller Synthesis arxiv: , K Zhou, J C Doyle, and K Glover Robust and Optimal Control

13 A Background on System Level Synthesis We begin by defining two function spaces which we use extensively throughout: RH = {M : C C n p Mz is rational, Mz is analytic on D c }, RH C, ρ = {M RH Mk Cρ k, k = 1, 2, } A1 A2 Note that we use SC, ρ to denote RH C, ρ in the main body of the text Recall that our main object of interest is the system x k+1 = Ax k + Bu k + w k, and our goal is to design a LTI feedback control policy u = Kx such that the resulting closed loop system is stable For a given K, we refer to the closed loop transfer functions from w x and w u as the system response Symbolically, we denote these maps as Φ x and Φ u Simple algebra shows that given K, these maps take on the form Φ x = zi A BK, Φ u = KzI A BK A3 We then have the following theorem parameterizing the set of such stable closed-loop transfer functions that are achievable by a stabilizing controller K Theorem A1 State-Feedback Parameterization 25 The following are true: The affine subspace defined by Φ zi A B x Φ u = I, Φ x, Φ u 1 z RH A4 parameterizes all system responses A3 from w to x, u, achievable by an internally stabilizing state-feedback controller K For any transfer matrices {Φ x, Φ u } satisfying A4, the controller K = Φ u Φ x stabilizing and achieves the desired system response A3 is internally If K stabilizes A, B, then the LQR cost of K on A, B can be written by Parseval s identity as T JA, B, K; σwi 2 1 := lim T T E x k Qx k + u k Ru k = σw 2 k=1 Q 1/2 0 Φx 2 0 R 1/2 A5 Φ u H 2 More generally, we will define JA, B, K; Σ to be the LQR cost when the process noise is driven by w iid N 0, Σ When we omit the last argument, we mean σw 2 = 1, ie JA, B, K = JA, B, K; I In 8, the authors use SLS to study how uncertainty in the true parameters A, B affect the LQR objective cost Our analysis relies on these tools, which we briefly describe below The starting point for the theory is a characterization of all robustly stabilizing controllers 13

14 Theorem A2 18 Suppose that the transfer matrices {Φ x, Φ u } 1 z RH satisfy Φ zi A B x = I + A6 Φ u Then the controller K = Φ u Φ x stabilizes the system described by A, B if and only if I + RH Furthermore, the resulting system response is given by x Φx = I + w A7 u Φ u This robustness result is used to derive a cost perturbation result for LQR Lemma A3 8 Let the controller K stabilize Â, B and Φ x, Φ u be its corresponding system response on system Â, B Then if K stabilizes A, B, it achieves the following LQR cost Q 1 Φx 2 0 JA, B, K = I + Φ 0 R 1 2 Φ A x B A8 u Φ u H2 Furthermore, letting := A Φ x B A9 Φ u a sufficient condition for K to stabilize A, B is that H < 1 An upper bound on H is given by, for any α 0, 1, εa H α Φ x, A10 H ε B 1 α Φ u where we assume that A Â 2 ε A and B B 2 ε B B Synthesis Results We first study the following infinite-dimensional synthesis problem 1 minimize γ 0,1 1 γ min Q 1 Φx 2 0 Φ x,φ u 0 R 1 2 Φ u H2 st zi Â B Φ x = I, Φx Φ u Φ u H Φ x 1 z RH C x, ρ, Φ u 1 z RH C u, ρ γ 2ε B1 We will conduct our analysis assuming that this infinite-dimensional problem is solvable Later on, we will show how to relax this problem to a finite-dimension one via FIR truncation, and show the minor modifications needed to the analysis for the guarantees to hold We now prove a sub-optimality guarantee on the solution to B1 which holds for certain choices of ε and the coefficients C x, ρ x and C u, ρ u This result also establishes an important technical consideration, which is when the problemmmm B1 is feasible 14

15 Theorem B1 Let J denote the minimal LQR cost achievable by any controller for the dynamical system with transition matrices A, B, and let K denote its optimal static feedback contoller Suppose that R A +B K RH C, ρ and that wlog ρ 1/e Suppose furthermore that ε is small enough to satisfy the following conditions: ε1 + K R A +B K H 1/5, ε1 + K C 1 ρ Let Â, B be any estimates of the transition matrices such that max{ A, B } ε Then, if C x, ρ and C u, ρ are set as, C x = O1C 1 ρ, C u = O1 K C 1 ρ, ρ = 1/4ρ + 3/4, we have that a the program B1 is feasible, b letting K denote an optimal solution to B1, the relative error in the LQR cost is JA, B, K 1 + 5ε1 + K R A +B K H 2 J, B2 and c if furthermore εc x + C u 21 ρ, the response { Φ x, Φ u } of K on the true system A, B satisfies O1C Φ x RH 1 ρ 2, 7/8 + 1/8ρ, O1 K C Φ u RH 1 ρ 2, 7/8 + 1/8ρ Proof The proof of a and b is nearly identical to that given in 8, which works by showing that Φ x = RÂ+ BK and Φ u = K RÂ+ BK is a feasible response which gives the desired suboptimality guarantee The only modification is that we need to find constants C x, C u, ρ for which RÂ+ BK 1 z RH C x, ρ and K RÂ+ BK 1 z RH C u, ρ We do this by writing RÂ+ BK = R A +B K I, = A + B K R A +B K By the definition of and our assumptions, we have that RH ε1 + K C, ρ, H < 1 This places us in a position to apply Lemma F3, from which we conclude that I RH O1, Avgρ, 1 Now applying Lemma F1 to R A +B K I, we conclude that O1C RÂ+ RH BK, 1/4ρ + 3/4 1 ρ 15

16 The claims of a and b now follows Now for the proof of c Let {Φ x, Φ u } be the solution to B1 We have that Φx Φx = I + Φ u Φ Φ, = A x B u Φ u We know that H < 1 by the constraints of the optimization problem B1 and furthermore, RH εc x + C u, ρ By assumption we have εc x + C u 2, from which we conclude using Lemma F3 that I + RH O1, Avgρ, 1 Furthermore, from Lemma F1, we conclude that Φ x I + Cx RH, 3/4 + 1/4ρ 1 ρ Φ u I + Cu RH, 3/4 + 1/4ρ 1 ρ The claim now follows by plugging in the values of C x, C u, and ρ B1 Suboptimality bounds for FIR truncated SLS Optimization problem B1 is convex but infinite dimensional, and as far as we are aware does not admit an efficient solution In Algorithm 1, we instead propose solving the following FIR approximation to problem B1: 1 minimize γ 0,1 1 γ st zi Â Q 1/2 0 Φx 0 R 1/2 Φ u B Φ x = I + 1 z F V, min Φ x,φ u,v Φ u H 2 2ε 1 C x ρ F +1, Φx γ Φ u H V C x ρ F +1, Φ x 1 z RHF C x, ρ, Φ u 1 z RHF C u, ρ B3 where here F denotes the FIR truncation length used This optimization problem can be posed as a finite dimensional semidefinite program see Section G3 Let KF denote the resulting controller We begin with a lemma identifying conditions under which optimization problem B3 is feasible to ease notation going forward, we let ζ := ε1 + K R A +B K H Lemma B2 Let the assumptions of Theorem B1 hold, and further assume that F 0 log2c x log1/ρ 1 Then optimization problem B3 is feasible for any F F 0 16

17 Proof We construct a feasible solution as follows Let Φ x = RÂ+ BK 1 : F, Φ u = K RÂ+ BK 1 : F, V = RÂ+ F + 1, and γ = 2 2ζ BK 1 ζ First, the proposed Φ x, Φ u are FIR of length F, and hence, using the same arguments as in the proof of Theorem B1, Φ x RH F C x, ρ and Φ u RH F C u, ρ It then also follows immediately that V = RÂ+ F + 1 C BK x ρ F +1 Note that the affine constraint is equivalent to zi Â B Φ x Φ u = I + 1 z F V Φ x t + 1 = ÂΦ xt + BΦ u t, Φ x 1 = I, B4 for 1 t < F We have by construction that the proposed Φ x and Φ u satisfy this constraint Further, the combination of the FIR constraints and the affine constraint B4 impose that Φ x F + 1 = ÂΦ xf + BΦ u F V = 0 Now notice that for the proposed Φ x, Φ u, we have that ÂΦ xf + BΦ u F = Â+ BK RÂ+ BK F = RÂ+ F + 1, where the last equality follows from the fact that RÂ+ t + 1 = Â + BK BK BK t It follows that Φ x F + 1 = 0, as desired It remains to prove that 2ε 1 C x ρ F +1 Φx Φ u 2 2ζ H 1 ζ < 1 The final inequality follows immediately from the assumption that ζ 1/5 Further, note that 2ε Φx 1 C x ρ F +1 2 Φ u 2ε RÂ+ BK K H RÂ+ BK 2 2ζ 1 ζ, H where the first inequality follows from the assumption on on F 0 and that the proposed Φ x is a truncation of RÂ+ BK and that the proposed Φ u is a truncation of K RÂ+ BK, and final inequality follows by applying the triangle inequality and the definition of ζ This proves the result Next, we use this to bound the suboptimality gap of the performance achieved by the controller implemented using the solutions of optimization problem B3 Lemma B3 Let the assumptions of Lemma B2 hold Fix any C J > 0, and further let F log1 + C J C x 1 log1/ρ Denote by Φ x F, Φ u F, V F, γf the optimal solution to optimization problem B3, and let KF = Φ u F Φ x F Then JA, B, KF 1 + C J O1ε1 + K R A +B K H 2 J B5 17

18 Proof Let := Φ A x F B I + 1 Φ u F z F V F Further note that using a similar argument to that in the proof of Lemma 42 of 8, one can verify that 2ε H Φx F 1 C x ρ F +1 γf, Φ u F H where we have exploited that Φ x F, Φ u F, V F, γf form a feasible solution to optimization problem B3 Then, repeated application of Theorem A2 tells us that the performance achieved by KF on the true system is given by Q 1 Φx 2 0 F JA, B, KF = I R 1 2 Φ u F z F V F I + H2 1 1 Q 1 Φx 2 0 F 1 C x ρ F +1, 1 γf 0 R 1 2 Φ u F H2 where the inequality follows from H γf < 1, and V F 1/2 by the assumption of F F 0 Denote by Φ x, Φ u, V, γ 0 the feasible solution constructed in the proof of Lemma B2 Then, C x ρ F +1 1 γf Q R 1 2 Φx F Φ u F H2 = C x ρ F C x ρ F C x ρ F +1 Q 1 Φx γ 0 0 R 1 2 Φ u J F 1 γ Â, B, K 0 JÂ, B, K 1 γ C x ρ F +1 1 γ 0 1 ζ J, where the first inequality follows from the optimality of Φ x F, Φ u F, V F, γf, the equality and second inequality from the fact that Φ x, Φ u are truncations of the response of K on Â, B to the first F time steps, and the final inequality by following similar arguments to the proof of Theorem 41 in 8 in applying Theorem A2 and noting that A RÂ+ BK B ζ < 1 K RÂ+ BK H We therefore have that JA, B, KF J 1 C x ρ F C J 1 γ 0 1 ζ γ 0 1 ζ J, where the last inequality follows from the assumptions on F stated in the Lemma Finally, by assumption ζ 1/5 < , from which it follows that 1 γ 0 1 ζ ζ, leading to the bound JA, B, KF 1 + C J ζ J H2 18

19 Squaring both sides proves the result The following Theorem is then immediate Theorem B4 Let J denote the minimal LQR cost achievable by any controller for A, B Let K denote the optimal controller and suppose that R A +B K RH C, ρ Fix a C J > 0, and suppose that F 0 and ε satisfy the assumptions of Lemmas B2 and B3 Let Â, B be any estimates of the transition matrices such that max{ A, B } ε Then, if C x, ρ and C u, ρ are set as in Lemma B2, we have that a the program B3 is feasible for any truncation length F F 0, b letting KF denote an optimal solution to B3 for truncation length F, the relative error in the LQR cost is JA, B, KF 1 + C J O1ε1 + K R A +B K H 2 J, B6 and c if furthermore εc x + C u O11 ρ 2, the response { Φ x, Φ u } of K on the true system A, B satisfies O1C Φ x RH 1 ρ 3, ρ, O1 K C Φ u RH 1 ρ 3, ρ Proof Claims a and b follow immediately from Lemmas B2 and B3 Now for the proof of c Let {Φ x F, Φ u F } be the solution to B3 Then as argued in the proof of Lemma B3, the response achieved on the true system A, B is given by Φx F Φ u F I + 1 z F V F I +, where is defined as in the proof of Lemma B3 We start by noting that Φ x F RH C x, ρ, and by the assumption on F F 0, it holds that z F V F RH 2, ρ 1/2 This allows us to apply Lemma F3 to conclude that I + z F V F RH O11 ρ 1/2, Avgρ 1/2, 1 Thus, applying Lemma F1 we conclude that Φ x F I + 1 z F V F O1Cx RH 1 ρ 1/2, AvgAvgρ1/2, 1, 1 A similar argument yields Φ u F I + 1 z F V F O1Cu RH 1 ρ 1/2, AvgAvgρ1/2, 1, 1 Now note that = A Φ x F + B Φ u F I + z F V F From the previous argument, we have that A Φ x F I + z F V F RH ε O1C x 1 ρ 1/2, AvgAvgρ1/2, 1, 1, B Φ u F I + z F V F RH ε O1C u 1 ρ 1/2, AvgAvgρ1/2, 1, 1, 19

20 from which it follows that RH ε O1C x + C u 1 ρ 1/2, AvgAvgρ 1/2, 1, 1 By the assumptions of the Theorem, we have that ε O1Cx+Cu 2, allowing us to apply Lemma F3 1 ρ 1/2 to conclude that I + RH O1, AvgAvgAvgρ 1/2, 1, 1, 1 Applying Lemma F1, we see that Φ x F I + z F V F I + O1Cx RH 1 ρ 1/2, AvgAvgAvgAvgρ1/2, 1, 1, 11 Φ u F I + z F V F I + O1Cu RH 1 ρ 1/2, AvgAvgAvgAvgρ1/2, 1, 1, 11 Finally, to simplify these bounds to those in the Theorem statement, notice first that for ρ 4, we have that 1 ρ 1/2 > 1 ρ 2 Then, we also have that AvgAvgAvgAvgρ 1/2, 1, 1, 11 = ρ1/2 = ρ /2 Finally, one can check that for ρ 4, we have that 1 4 ρ / ρ, leading to the bound ρ / ρ ρ We note that these constants are by no means optimized C Estimation Recall that Algorithm 1 proceeds in epochs and that we denote by x i t and u i t the state and input at time t during epoch i, respectively The i-th epoch has length T i Note that x i T i, the last state of epoch i, is equal to x i+1 0, the first state of epoch i + 1 At the end of each epoch our method estimates the parameters A, B from the trajectory observed during that epoch, ie T Â, B i 1 arg min A,B 2 xi t+1 Axi t Bu i t 2 2 C1 The goal of this section is to offer high probability confidence bounds on the estimation error of C1 For the rest of the section we suppress the dependence on the epoch index i because we prove a statistical rate for a fixed epoch Algorithm 1 generates the inputs u t using a feedback controller K which stabilizes the true system A, B Let {Φ x, Φ u } denote the response of K on the true system A, B, and suppose 20

21 that Φ x 1 z RH C x, ρ and Φ u 1 z RH iid C u, ρ More precisely, if w t N 0, σwi 2 p is the process iid noise at time t and η t N 0, σηi 2 p is the input noise added at time t, then we can write t x t = Φ x t + 1x 0 + Φ x t kb η k + w k k=0 t u t = η t + Φ u t + 1x 0 + Φ u t kb η k + w k k=0 C2 C3 For the statistical analysis it is useful to consider the stochastic process z t = x t, u t Also, we denote the filtration F t = σx 0, η 0, w 0, η t, w t, η t It is clear that the process {z t } t 0 is {F t } t 0 -adapted Throughout this section we assume that C u, C x 1 and denote C 2 K := nc2 x + pc 2 u C1 Estimation after one epoch Throughout this section we assume that σ η σ w This condition is not needed for achieving the necessary statistical rate of estimation of A, B, but it aids in simplifying several algebraic quantities Proposition C1 Let x 0 R n be any initial state, let σ η σ w, and assume that a trajectory {x t, u t } T is observed Furthermore, suppose the inputs u t R p are generated by a feedback controller K which stabilizes and achieves a response {Φ x, Φ u } on A, B with Φ x 1 z RH C x, ρ and Φ u 1 z RH C u, ρ Then, the error of the OLS estimator Â, B from Eq C1 satisfies with probability 1 δ the guarantee { } max Â A, B B σ wc u n + p log 1 + pc u + σ w ρc u C K σ η T δ σ η δ1 ρ B + x 0 2, σ w T as long as T n + p log 1 + pc2 u + σ2 w ρ 2 CuC 2 2 K δ ση 2 δ1 ρ B 2 + x σwt 2 C4 The proof of this result follows from a result by Simchowitz et al 23 on the estimation of linear response time-series We present that result in the context of our problem Let M = A, B, and recall that z t = x t, y t Then, the OLS estimator C1 can be written in the form 1 M arg min M 2 x t+1 Mz t 2 2 C5 The process {z t } t 0 is said to satisfy the k, ν, β-block martingale small-ball BMSB condition if for any j 0 and v R n+p, one has that 1 k k P v, z j+i ν β almost surely i=1 This condition is used for characterizing the size of the minimum eigenvalue of the covariance matrix T z tz t A larger ν guarantees a larger lower bound of the minimum eigenvalue In the context of our problem the result by Simchowitz et al 23 translates as follows 21

22 Theorem C2 Simchowitz et al 23 Fix ɛ, δ 0, 1 For every T, k, ν, and β such that {z t } t 0 satisfies the k, ν, β-bmsb and T n + p T k β 2 log 1 + TrEz tzt k T/k β 2 ν 2, δ the estimate M defined in Eq C5 satisfies the following statistical rate P M M 2 > O1σ w n + p T βν k T/k log 1 + TrEz tzt k T/k β 2 ν 2 δ δ Therefore, in order to apply this result we need to find k, ν, and β such that {z t } t 0 satisfies the k, ν, β-bmsb condition, and we also have to upper bound the trace of the covariance of z t The next two lemmas address these two issues Lemma C3 Let x 0 be any initial state in R n and let {u t } t 0 be the sequence of inputs generated according to C3, and assume σ η σ w Then, the process z t = x t, u t satisfies the Proof For all t 1, denote Therefore, we have σ η 1,, 3 2C u 20 ξ t = u t η t Φ u 1w t BMSB condition t 2 = Φ u t + 1x 0 + Φ u t kb η k + w k + Φ u 1B η t xt+1 u t+1 k=0 A x = t + B u t In 0 wt + Φ u 1 ξ t+1 I p η t+1, and hence xt+1 u t+1 F t N A x t + B u t σ 2, w I n σwφ 2 u 1 σwφ 2 u 1 σwφ 2 u 1Φ u 1 + σηi 2 p ξ t+1 Denote by µ z,t and Σ z the mean and covariance of this multivariate normal distribution Recall that we denote z t = x t, u t Let v R n+p and then v, z t N v, µ z,t, v Σ z v Therefore, P v, z t λ min Σ z P v, z t v Σ z v P v, z t µ z,t v Σ z v 3/10, where the last two inequalities follow because for any µ, σ 2 R and ω N 0, σ 2 we have P µ + ω σ P ω σ 3/10 22

23 Since Φ u 1 z RH C u, ρ we have Φ u 1 C u Then, by a simple argument based on a Schur complement detailed in Lemma F6 it follows that The conclusion follows since C u 1 1 λ min Σ z ση 2 min 2, σw 2 2σwC 2 u 2 + ση 2 Lemma C4 Let σ η σ w Then, the process z t = x t, u t satisfies Tr Ez t zt Proof Now, note that σηpt 2 + σw 2 ρ 2 CK 2 T 1 ρ B 2 + x σwt 2 Ez t zt Φx t + 1 = x Φ u t x Φx t Φ u t σηi 2 p t Φx t k + σ 2 Φ u t k ηb B + σwi 2 Φx t k n Φ u t k k=0 Since for all j 1 we have Φ x j C x ρ j and Φ u j C u ρ j, we obtain t Tr Ez t zt pση 2 + ncx 2 + pcu 2 ρ 2t+2 x σw 2 + ση B 2 2 ρ 2k Therefore, we get that Tr Ez t zt pσηt 2 + ρ2 T 1 ρ 2 nc2 x + pcuσ 2 w 2 + ση B ρ2 1 ρ 2 nc2 x + pcu x , and the conclusion follows by simple algebra Proposition C1 follows from Theorem C2, Lemma C3, Lemma C4, and simple algebra C2 Stitching the epochs together We start by bounding with high probability the size of the initial states of the epochs Recall that epoch i has length T i and that we denote by x i T i the last state of the epoch i, which is equal to the first state x i+1 0 of the epoch i + 1 For simplicity we assume that x 0 0 = 0, an assumption that is not restrictive in any way Lemma C5 Fix δ 0, 1, r > 0, and an epoch i Assume that for all k i the epoch length T k is large enough so that C x ρ T k ρ r Then, for any t 0 we have P x i σ w n + t C xρ1 + B 1 ρ r 1 ρ 2 exp t2 2 k=1 23

Lecture 20: Linear Dynamics and LQG

CSE599i: Online and Adaptive Machine Learning Winter 2018 Lecturer: Kevin Jamieson Lecture 20: Linear Dynamics and LQG Scribes: Atinuke Ademola-Idowu, Yuanyuan Shi Disclaimer: These notes have not been