Use of variance estimation in the multi-armed bandit problem

Jean-Yves Audibert; Rémi Munos; Csaba Szepesvari

Autre Publication Année : 2006

Use of variance estimation in the multi-armed bandit problem

(1) , (2) , (3)

1
2
3

Jean-Yves Audibert

Fonction : Auteur
PersonId : 931557

Centre d'Enseignement et de Recherche en Mathématiques et Calcul Scientifique

Rémi Munos

Fonction : Auteur
PersonId : 836863

Sequential Learning

Csaba Szepesvari

Fonction : Auteur

Computer and Automation Research Institute [Budapest]

Résumé

An important aspect of most decision making problems concerns the appropriate balance between exploitation (acting optimally according to the partial knowledge acquired so far) and exploration of the environment (acting sub-optimally in order to refine the current knowledge and improve future decisions). A typical example of this so-called exploration versus exploitation dilemma is the multi-armed bandit problem, for which many strategies have been developed. Here we investigate policies based the choice of the arm having the highest upper-confidence bound, where the bound takes into account the empirical variance of the different arms. Such an algorithm was found earlier to outperform its peers in a series of numerical experiments. The main contribution of this paper is the theoretical investigation of this algorithm. Our contribution here is twofold. First, we prove that with probability at least $1-\beta$, the regret after $n$ plays of a variant of the UCB algorithm (called $\beta$-UCB) is upper-bounded by a constant, that scales linearly with $\log(1/\beta)$, but which is independent from $n$. We also analyse a variant which is closer to the algorithm suggested earlier. We prove a logarithmic bound on the expected regret of this algorithm and argue that the bound scales favourably with the variance of the suboptimal arms.

Domaines

Apprentissage [cs.LG]

Fichier principal

ucbtuned.pdf (124.66 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Rémi Munos : Connectez-vous pour contacter le contributeur

https://inria.hal.science/inria-00203496

Soumis le : jeudi 10 janvier 2008-12:11:52

Dernière modification le : vendredi 24 mars 2023-14:52:49

Archivage à long terme le : mardi 13 avril 2010-16:56:00

Dates et versions

inria-00203496 , version 1 (10-01-2008)

Identifiants

HAL Id : inria-00203496 , version 1

Citer

Jean-Yves Audibert, Rémi Munos, Csaba Szepesvari. Use of variance estimation in the multi-armed bandit problem. 2006. ⟨inria-00203496⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

ENPC UNIV-LILLE3 CNRS INRIA CERMICS PARISTECH LAGIS INRIA2

292 Consultations

283 Téléchargements

Use of variance estimation in the multi-armed bandit problem

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager