APRIL: Active Preference-learning based Reinforcement Learning

Riad Akrour; Marc Schoenauer; Michèle Sebag

Communication Dans Un Congrès Année : 2012

APRIL: Active Preference-learning based Reinforcement Learning

(1, 2) , (1, 2) , (2)

1
2

Riad Akrour

Fonction : Auteur
PersonId : 910562

Machine Learning and Optimisation

Laboratoire de Recherche en Informatique

Marc Schoenauer

Fonction : Auteur
PersonId : 739309
IdHAL : evomarc
ORCID : 0000-0003-1450-6830
IdRef : 057775575

Machine Learning and Optimisation

Laboratoire de Recherche en Informatique

Michèle Sebag

Fonction : Auteur
PersonId : 836537

Laboratoire de Recherche en Informatique

Résumé

This paper focuses on reinforcement learning (RL) with limited prior knowledge. In the domain of swarm robotics for instance, the expert can hardly design a reward function or demonstrate the target behavior, forbidding the use of both standard RL and inverse reinforcement learning. Although with a limited expertise, the human expert is still often able to emit preferences and rank the agent demonstrations. Earlier work has presented an iterative preference-based RL framework: expert preferences are exploited to learn an approximate policy return, thus enabling the agent to achieve direct policy search. Iteratively, the agent selects a new candidate policy and demonstrates it; the expert ranks the new demonstration comparatively to the previous best one; the expert's ranking feedback enables the agent to refine the approximate policy return, and the process is iterated. In this paper, preference-based reinforcement learning is combined with active ranking in order to decrease the number of ranking queries to the expert needed to yield a satisfactory policy. Experiments on the mountain car and the cancer treatment testbeds witness that a couple of dozen rankings enable to learn a competent policy.

Mots clés

reinforcement learning preference learning interactive optimization robotics

Domaines

Intelligence artificielle [cs.AI]

Fichier principal

April_camera.pdf (330.79 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Marc Schoenauer : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-00722744

Soumis le : vendredi 3 août 2012-17:50:20

Dernière modification le : lundi 12 février 2024-09:48:04

Archivage à long terme le : vendredi 16 décembre 2016-05:15:39

Dates et versions

hal-00722744 , version 1 (03-08-2012)

Identifiants

HAL Id : hal-00722744 , version 1
ARXIV : 1208.0984

Citer

Riad Akrour, Marc Schoenauer, Michèle Sebag. APRIL: Active Preference-learning based Reinforcement Learning. ECML PKDD 2012, Sep 2012, Bristol, United Kingdom. pp.116-131. ⟨hal-00722744⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

EC-PARIS UNIV-RENNES1 CNRS INRIA IRISA UMR8623 INRIA2 LRI-AO UR1-MATH-STIC UNIV-PARIS-SACLAY UR1-UFR-ISTIC UNIV-RENNES UR1-MATH-NUM

284 Consultations

255 Téléchargements

APRIL: Active Preference-learning based Reinforcement Learning

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager