Factorial scaled hidden Markov model for polyphonic audio representation and source separation - Inria - Institut national de recherche en sciences et technologies du numérique Accéder directement au contenu
Communication Dans Un Congrès Année : 2009

Factorial scaled hidden Markov model for polyphonic audio representation and source separation

Résumé

We present a new probabilistic model for polyphonic audio termed Factorial Scaled Hidden Markov Model (FS-HMM), which generalizes several existing models, notably the Gaussian scaled mixture model and the Itakura-Saito Nonnegative Matrix Factorization (NMF) model. We describe two expectation-maximization (EM) algorithms for maximum likelihood estimation, which differ by the choice of complete data set. The second EM algorithm, based on a reduced complete data set and multiplicative updates inspired from NMF methodology, exhibits much faster convergence. We consider the FS-HMM in different configurations for the difficult problem of speech / music separation from a single channel and report satisfying results.
Fichier principal
Vignette du fichier
OzerovFevotteCharbit_WASPAA09.pdf (171.52 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

inria-00553336 , version 1 (07-01-2011)

Identifiants

  • HAL Id : inria-00553336 , version 1

Citer

Alexey Ozerov, Cédric Févotte, Maurice Charbit. Factorial scaled hidden Markov model for polyphonic audio representation and source separation. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA'09), Oct 2009, Mohonk, NY, United States. ⟨inria-00553336⟩
138 Consultations
357 Téléchargements

Partager

Gmail Facebook X LinkedIn More