Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles

Yann Ollivier; Ludovic Arnold; Anne Auger; Nikolaus Hansen

Article Dans Une Revue Journal of Machine Learning Research Année : 2017

Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles

(1, 2, 3) , (3, 2) , (4, 5, 3) , (4, 5, 3)

1
2
3
4
5

Yann Ollivier

Fonction : Auteur
PersonId : 883809

TAckling the Underspecified

Laboratoire de Recherche en Informatique

Machine Learning and Optimisation

Ludovic Arnold

Fonction : Auteur

Machine Learning and Optimisation

Laboratoire de Recherche en Informatique

Anne Auger

Fonction : Auteur
PersonId : 751513
IdHAL : anne-auger

Centre de Mathématiques Appliquées - Ecole Polytechnique

Randomized Optimisation

Machine Learning and Optimisation

Nikolaus Hansen

Fonction : Auteur
PersonId : 1943
IdHAL : nikolaus-hansen
ORCID : 0000-0001-7788-4906
IdRef : 154974463

Centre de Mathématiques Appliquées - Ecole Polytechnique

Randomized Optimisation

Machine Learning and Optimisation

Résumé

We present a canonical way to turn any smooth parametric family of probability distributions on an arbitrary search space X into a continuous-time black-box optimization method on X, the information-geometric optimization (IGO) method. Invariance as a major design principle keeps the number of arbitrary choices to a minimum. The resulting IGO flow is the flow of an ordinary differential equation conducting the natural gradient ascent of an adaptive, time-dependent transformation of the objective function. It makes no particular assumptions on the objective function to be optimized. The IGO method produces explicit IGO algorithms through time discretization. It naturally recovers versions of known algorithms and offers a systematic way to derive new ones. In continuous search spaces, IGO algorithms take a form related to natural evolution strategies (NES). The cross-entropy method is recovered in a particular case with a large time step, and can be extended into a smoothed, parametrization-independent maximum likelihood update (IGO-ML). When applied to the family of Gaussian distributions on R^d, the IGO framework recovers a version of the well-known CMA-ES algorithm and of xNES. For the family of Bernoulli distributions on {0, 1}^d, we recover the seminal PBIL algorithm and cGA. For the distributions of restricted Boltzmann machines, we naturally obtain a novel algorithm for discrete optimization on {0, 1}^d. All these algorithms are natural instances of, and unified under, the single information-geometric optimization framework. The IGO method achieves, thanks to its intrinsic formulation, maximal invariance properties: invariance under reparametrization of the search space X, under a change of parameters of the probability distribution, and under increasing transformation of the function to be optimized. The latter is achieved through an adaptive, quantile-based formulation of the objective. Theoretical considerations strongly suggest that IGO algorithms are essentially characterized by a minimal change of the distribution over time. Therefore they have minimal loss in diversity through the course of optimization, provided the initial diversity is high. First experiments using restricted Boltzmann machines confirm this insight. As a simple consequence, IGO seems to provide, from information theory, an elegant way to simultaneously explore several valleys of a fitness landscape in a single run.

Mots clés

black-box optimization stochastic optimization randomized optimization natural gradient invariance evolution strategy information-geometric optimization

Domaines

Apprentissage [cs.LG] Réseau de neurones [cs.NE] Analyse numérique [cs.NA]

Fichier principal

14-467.pdf (914.78 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Nikolaus Hansen : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-01515898

Soumis le : vendredi 28 avril 2017-11:37:49

Dernière modification le : lundi 12 février 2024-09:48:04

Archivage à long terme le : samedi 29 juillet 2017-13:01:21

Dates et versions

hal-01515898 , version 1 (28-04-2017)

Identifiants

HAL Id : hal-01515898 , version 1

Citer

Yann Ollivier, Ludovic Arnold, Anne Auger, Nikolaus Hansen. Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles. Journal of Machine Learning Research, 2017, 18 (18), pp.1-65. ⟨hal-01515898⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

X CNRS INRIA X-CMAP X-DEP-MATHA CMAP UMR8623 CENTRALESUPELEC INRIA2 LRI-AO UNIV-PARIS-SACLAY LISN GS-ENGINEERING GS-COMPUTER-SCIENCE LISN-AO

677 Consultations

398 Téléchargements

Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager