Blind Audiovisual Source Separation Based on Sparse Redundant Representations

Anna Llagostera Casanovas; Monaci Gianluca; Pierre Vandergheynst; Rémi Gribonval

doi:10.1109/TMM.2010.2050650

Article Dans Une Revue IEEE Transactions on Multimedia Année : 2010

Blind Audiovisual Source Separation Based on Sparse Redundant Representations

(1) , (1, 2) , (1) , (3)

1
2
3

Anna Llagostera Casanovas

Fonction : Auteur
PersonId : 884023

LTS2 - EPFL

Monaci Gianluca

Fonction : Auteur
PersonId : 884024

LTS2 - EPFL

Philips Research [Nederlands]

Pierre Vandergheynst

Fonction : Auteur
PersonId : 839985

LTS2 - EPFL

Rémi Gribonval

Fonction : Auteur
PersonId : 1255
IdHAL : remi-gribonval
ORCID : 0000-0002-9450-8125
IdRef : 113181590

Speech and sound data modeling and processing

Résumé

In this paper, we propose a novel method which is able to detect and separate audiovisual sources present in a scene. Our method exploits the correlation between the video signal captured with a camera and a synchronously recorded one-microphone audio track. In a first stage, audio and video modalities are decomposed into relevant basic structures using redundant representations. Next, synchrony between relevant events in audio and video modalities is quantified. Based on this co-occurrence measure, audiovisual sources are counted and located in the image using a robust clustering algorithm that groups video structures exhibiting strong correlations with the audio. Next periods where each source is active alone are determined and used to build spectral Gaussian mixture models (GMMs) characterizing the sources acoustic behavior. Finally, these models are used to separate the audio signal in periods during which several sources are mixed. The proposed approach has been extensively tested on synthetic and natural sequences composed of speakers and music instruments. Results show that the proposed method is able to successfully detect, localize, separate, and reconstruct present audiovisual sources.

Mots clés

Audiovisual processing blind source separation Gaussian mixture models sparse signal representation. sparse signal representation

Domaines

Théorie de l'information [cs.IT] Théorie de l'information et codage [math.IT]

Fichier principal

2010_IEEE_TMM_LlagosteraEtAl_BAVSS.pdf (5.03 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Rémi Gribonval : Connectez-vous pour contacter le contributeur

https://inria.hal.science/inria-00541412

Soumis le : jeudi 27 janvier 2011-21:58:42

Dernière modification le : mardi 16 avril 2024-11:11:45

Archivage à long terme le : jeudi 28 avril 2011-02:27:58

Dates et versions

inria-00541412 , version 1 (27-01-2011)

Identifiants

HAL Id : inria-00541412 , version 1
DOI : 10.1109/TMM.2010.2050650

Citer

Anna Llagostera Casanovas, Monaci Gianluca, Pierre Vandergheynst, Rémi Gribonval. Blind Audiovisual Source Separation Based on Sparse Redundant Representations. IEEE Transactions on Multimedia, 2010, 12 (5), pp.358 -- 371. ⟨10.1109/TMM.2010.2050650⟩. ⟨inria-00541412⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

EC-PARIS UNIV-RENNES1 CNRS INRIA INSA-RENNES IRISA IRISA-D5 INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES INSA-GROUPE UR1-MATH-NUM

302 Consultations

374 Téléchargements

Blind Audiovisual Source Separation Based on Sparse Redundant Representations

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager