SSC : Statistical Subspace Clustering

Subspace clustering is an extension of traditional clustering that seeks to find clusters in different subspaces within a dataset. This is a particularly important challenge with high dimensional data where the curse of dimensionality occurs. It has also the benefit of providing smaller descriptions of the clusters found. Existing methods only consider numerical databases and do not propose any method for clusters visualization. Besides, they require some input parameters difficult to set for the user. The aim of this paper is to propose a new subspace clustering algorithm, able to tackle databases that may contain continuous as well as discrete attributes, requiring as few user parameters as possible, and producing an interpretable output. We present a method based on the use of the well-known EM algorithm on a probabilistic model designed under some specific hypotheses, allowing us to present the result as a set of rules, each one defined with as few relevant dimensions as possible. Experiments, conducted on artificial as well as real databases, show that our algorithm gives robust results, in terms of classification and interpretability of the output.

Domaines

Langage de programmation [cs.PL]

Fichier principal

MLDM05.pdf (135.39 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Isabelle Tellier : Connectez-vous pour contacter le contributeur

https://inria.hal.science/inria-00536697

Soumis le : mardi 16 novembre 2010-17:32:22

Dernière modification le : mercredi 19 avril 2023-04:22:44

Archivage à long terme le : jeudi 17 février 2011-03:07:31

Dates et versions

inria-00536697 , version 1 (16-11-2010)

Identifiants

HAL Id : inria-00536697 , version 1

Citer

Laurent Candillier, Isabelle Tellier, Fabien Torre, Olivier Bousquet. SSC : Statistical Subspace Clustering. 4th International Conference on Machine Learning and Data Mining in Pattern Recognition, 2005, Leipzig, Georgia. pp.100--109. ⟨inria-00536697⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-LILLE3 CNRS INRIA MOSTRARE INRIA2

139 Consultations

396 Téléchargements