Blind Audiovisual Source Separation Based on Sparse Redundant Representations

Anna Llagostera Casanovas 1 Monaci Gianluca 1, 2 Pierre Vandergheynst 1 Rémi Gribonval 3
3 METISS - Speech and sound data modeling and processing
IRISA - Institut de Recherche en Informatique et Systèmes Aléatoires, Inria Rennes – Bretagne Atlantique
Abstract : In this paper, we propose a novel method which is able to detect and separate audiovisual sources present in a scene. Our method exploits the correlation between the video signal captured with a camera and a synchronously recorded one-microphone audio track. In a first stage, audio and video modalities are decomposed into relevant basic structures using redundant representations. Next, synchrony between relevant events in audio and video modalities is quantified. Based on this co-occurrence measure, audiovisual sources are counted and located in the image using a robust clustering algorithm that groups video structures exhibiting strong correlations with the audio. Next periods where each source is active alone are determined and used to build spectral Gaussian mixture models (GMMs) characterizing the sources acoustic behavior. Finally, these models are used to separate the audio signal in periods during which several sources are mixed. The proposed approach has been extensively tested on synthetic and natural sequences composed of speakers and music instruments. Results show that the proposed method is able to successfully detect, localize, separate, and reconstruct present audiovisual sources.
Type de document :
Article dans une revue
IEEE Transactions on Multimedia, Institute of Electrical and Electronics Engineers, 2010, 12 (5), pp.358 -- 371. 〈10.1109/TMM.2010.2050650〉
Liste complète des métadonnées

Littérature citée [29 références]  Voir  Masquer  Télécharger

https://hal.inria.fr/inria-00541412
Contributeur : Rémi Gribonval <>
Soumis le : jeudi 27 janvier 2011 - 21:58:42
Dernière modification le : jeudi 11 janvier 2018 - 06:20:09
Document(s) archivé(s) le : jeudi 28 avril 2011 - 02:27:58

Fichier

2010_IEEE_TMM_LlagosteraEtAl_B...
Fichiers produits par l'(les) auteur(s)

Identifiants

Collections

Citation

Anna Llagostera Casanovas, Monaci Gianluca, Pierre Vandergheynst, Rémi Gribonval. Blind Audiovisual Source Separation Based on Sparse Redundant Representations. IEEE Transactions on Multimedia, Institute of Electrical and Electronics Engineers, 2010, 12 (5), pp.358 -- 371. 〈10.1109/TMM.2010.2050650〉. 〈inria-00541412〉

Partager

Métriques

Consultations de la notice

398

Téléchargements de fichiers

234