Making Deep Q-learning methods robust to time discretization

Corentin Tallec; Léonard Blier; Yann Ollivier

Communication Dans Un Congrès Année : 2019

Making Deep Q-learning methods robust to time discretization

(1) , (2, 1) , (2, 1)

1
2

Corentin Tallec

Fonction : Auteur

TAckling the Underspecified

Léonard Blier

Fonction : Auteur

Facebook AI Research [Paris]

TAckling the Underspecified

Yann Ollivier

Fonction : Auteur
PersonId : 883809

Facebook AI Research [Paris]

TAckling the Underspecified

Résumé

Despite remarkable successes, Deep Reinforcement Learning (DRL) is not robust to hyperparameterization, implementation details, or small environment changes (Henderson et al. 2017, Zhang et al. 2018). Overcoming such sensitivity is key to making DRL applicable to real world problems. In this paper, we identify sensitivity to time discretization in near continuous-time environments as a critical factor; this covers, e.g., changing the number of frames per second, or the action frequency of the controller. Empirically, we find that Q-learning-based approaches such as Deep Q- learning (Mnih et al., 2015) and Deep Deterministic Policy Gradient (Lillicrap et al., 2015) collapse with small time steps. Formally, we prove that Q-learning does not exist in continuous time. We detail a principled way to build an off-policy RL algorithm that yields similar performances over a wide range of time discretizations, and confirm this robustness empirically.

Mots clés

Time discretization Robustness Reinforcement Learning Deep Learning

Domaines

Intelligence artificielle [cs.AI]

Marc Schoenauer : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-02435523

Soumis le : vendredi 10 janvier 2020-20:11:27

Dernière modification le : mardi 13 février 2024-03:38:17

Dates et versions

hal-02435523 , version 1 (10-01-2020)

Identifiants

HAL Id : hal-02435523 , version 1
ARXIV : 1901.09732

Citer

Corentin Tallec, Léonard Blier, Yann Ollivier. Making Deep Q-learning methods robust to time discretization. ICML 2019 - Thirty-sixth International Conference on Machine Learning, Jun 2019, Long Beach, United States. pp.6096-6104. ⟨hal-02435523⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS INRIA UMR8623 CENTRALESUPELEC INRIA2 LRI-AO UNIV-PARIS-SACLAY LISN GS-ENGINEERING GS-COMPUTER-SCIENCE LISN-AO

76 Consultations

0 Téléchargements

Making Deep Q-learning methods robust to time discretization

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager