Skip to Main content Skip to Navigation
Journal articles

A consolidated perspective on multi-microphone speech enhancement and source separation

Abstract : Speech enhancement and separation are core problems in audio signal processing, with commercial applications in devices as diverse as mobile phones, conference call systems, hands-free systems, or hearing aids. In addition, they are crucial pre-processing steps for noise-robust automatic speech and speaker recognition. Many devices now have two to eight microphones. The enhancement and separation capabilities offered by these multichannel interfaces are usually greater than those of single-channel interfaces. Research in speech enhancement and separation has followed two convergent paths, starting with microphone array processing and blind source separation, respectively. These communities are now strongly interrelated and routinely borrow ideas from each other. Yet, a comprehensive overview of the common foundations and the differences between these approaches is lacking at present. In this article, we propose to fill this gap by analyzing a large number of established and recent techniques according to four transverse axes: a) the acoustic impulse response model, b) the spatial filter design criterion, c) the parameter estimation algorithm, and d) optional postfiltering. We conclude this overview paper by providing a list of software and data resources and by discussing perspectives and future trends in the field.
Document type :
Journal articles
Complete list of metadata
Contributor : Emmanuel Vincent Connect in order to contact the contributor
Submitted on : Saturday, March 4, 2017 - 10:57:43 PM
Last modification on : Wednesday, November 3, 2021 - 7:09:38 AM
Long-term archiving on: : Tuesday, June 6, 2017 - 12:07:06 PM


Files produced by the author(s)




Sharon Gannot, Emmanuel Vincent, Shmulik Markovich-Golan, Alexey Ozerov. A consolidated perspective on multi-microphone speech enhancement and source separation. IEEE/ACM Transactions on Audio, Speech and Language Processing, Institute of Electrical and Electronics Engineers, 2017, 25 (4), pp.692-730. ⟨10.1109/TASLP.2016.2647702⟩. ⟨hal-01414179v2⟩



Les métriques sont temporairement indisponibles