Approximate Modified Policy Iteration

Bruno Scherrer 1 Victor Gabillon 2 Mohammad Ghavamzadeh 2 Matthieu Geist 3, 4
1 MAIA - Autonomous intelligent machine
Inria Nancy - Grand Est, LORIA - AIS - Department of Complex Systems, Artificial Intelligence & Robotics
2 SEQUEL - Sequential Learning
LIFL - Laboratoire d'Informatique Fondamentale de Lille, Inria Lille - Nord Europe, LAGIS - Laboratoire d'Automatique, Génie Informatique et Signal
Abstract : Modified policy iteration (MPI) is a dynamic programming (DP) algorithm that contains the two celebrated policy and value iteration methods. Despite its generality, MPI has not been thoroughly studied, especially its approximation form which is used when the state and/or action spaces are large or infinite. In this paper, we propose three implementations of approximate MPI (AMPI) that are extensions of well-known approximate DP algorithms: fitted-value iteration, fitted-Q iteration, and classification-based policy iteration. We provide error propagation analyses that unify those for approximate policy and value iteration. On the last classification-based implementation, we develop a finite-sample analysis that shows that MPI's main parameter allows to control the balance between the estimation error of the classifier and the overall value function approximation.
Type de document :
Rapport
[Research Report] 2012
Liste complète des métadonnées

Littérature citée [16 références]  Voir  Masquer  Télécharger

https://hal.inria.fr/hal-00697169
Contributeur : Bruno Scherrer <>
Soumis le : mercredi 16 mai 2012 - 17:02:59
Dernière modification le : jeudi 5 avril 2018 - 12:30:11
Document(s) archivé(s) le : vendredi 31 mars 2017 - 08:30:50

Fichiers

article.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : hal-00697169, version 2
  • ARXIV : 1205.3054

Citation

Bruno Scherrer, Victor Gabillon, Mohammad Ghavamzadeh, Matthieu Geist. Approximate Modified Policy Iteration. [Research Report] 2012. 〈hal-00697169v2〉

Partager

Métriques

Consultations de la notice

660

Téléchargements de fichiers

226