Policy iteration for perfect information stochastic mean payoff games with bounded first return times is strongly polynomial

Marianne Akian; Stéphane Gaubert

Pré-Publication, Document De Travail Année : 2013

Policy iteration for perfect information stochastic mean payoff games with bounded first return times is strongly polynomial

(1, 2) , (1, 2)

1
2

Marianne Akian

Fonction : Auteur
PersonId : 830429

Centre de Mathématiques Appliquées - Ecole Polytechnique

Max-plus algebras and mathematics of decision

Stéphane Gaubert

Fonction : Auteur
PersonId : 1887
IdHAL : stephane-gaubert
IdRef : 104895306

Centre de Mathématiques Appliquées - Ecole Polytechnique

Max-plus algebras and mathematics of decision

Résumé

Recent results of Ye and Hansen, Miltersen and Zwick show that policy iteration for one or two player (perfect information) zero-sum stochastic games, restricted to instances with a fixed discount rate, is strongly polynomial. We show that policy iteration for mean-payoff zero-sum stochastic games is also strongly polynomial when restricted to instances with bounded first mean return time to a given state. The proof is based on methods of nonlinear Perron-Frobenius theory, allowing us to reduce the mean-payoff problem to a discounted problem with state dependent discount rate. Our analysis also shows that policy iteration remains strongly polynomial for discounted problems in which the discount rate can be state dependent (and even negative) at certain states, provided that the spectral radii of the nonnegative matrices associated to all strategies are bounded from above by a fixed constant strictly less than 1.

Domaines

Optimisation et contrôle [math.OC]

Marianne Akian : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-00881207

Soumis le : jeudi 7 novembre 2013-17:14:04

Dernière modification le : lundi 22 avril 2024-14:01:57

Dates et versions

hal-00881207 , version 1 (07-11-2013)

Identifiants

HAL Id : hal-00881207 , version 1
ARXIV : 1310.4953

Citer

Marianne Akian, Stéphane Gaubert. Policy iteration for perfect information stochastic mean payoff games with bounded first return times is strongly polynomial. 2013. ⟨hal-00881207⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

X CNRS INRIA X-CMAP X-DEP-MATHA CMAP INRIA2 TDS-MACS

301 Consultations

0 Téléchargements

Policy iteration for perfect information stochastic mean payoff games with bounded first return times is strongly polynomial

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager