Skip to Main content Skip to Navigation

Rate of Convergence and Error Bounds for LSTD($\lambda$)

Manel Tagorti 1 Bruno Scherrer 1 
1 MAIA - Autonomous intelligent machine
Inria Nancy - Grand Est, LORIA - AIS - Department of Complex Systems, Artificial Intelligence & Robotics
Abstract : We consider LSTD($\lambda$), the least-squares temporal-difference algorithm with eligibility traces algorithm proposed by Boyan (2002). It computes a linear approximation of the value function of a fixed policy in a large Markov Decision Process. Under a $\beta$-mixing assumption, we derive, for any value of $\lambda \in (0,1)$, a high-probability estimate of the rate of convergence of this algorithm to its limit. We deduce a high-probability bound on the error of this algorithm, that extends (and slightly improves) that derived by Lazaric et al. (2012) in the specific case where $\lambda=0$. In particular, our analysis sheds some light on the choice of $\lambda$ with respect to the quality of the chosen linear space and the number of samples, that complies with simulations.
Complete list of metadata

Cited literature [12 references]  Display  Hide  Download
Contributor : Bruno Scherrer Connect in order to contact the contributor
Submitted on : Tuesday, May 13, 2014 - 3:49:54 PM
Last modification on : Saturday, June 25, 2022 - 7:39:44 PM
Long-term archiving on: : Monday, April 10, 2017 - 10:16:10 PM


Files produced by the author(s)


  • HAL Id : hal-00990525, version 1
  • ARXIV : 1405.3229


Manel Tagorti, Bruno Scherrer. Rate of Convergence and Error Bounds for LSTD($\lambda$). [Research Report] 2014. ⟨hal-00990525⟩



Record views


Files downloads