Robust ASR using neural network based speech enhancement and feature simulation

Sunit Sivasankaran 1 Aditya Arie Nugraha 1 Emmanuel Vincent 1 Juan Andrés Morales Cordovilla 1 Siddharth Dalmia 1 Irina Illina 2, 1 Antoine Liutkus 2, 1
1 MULTISPEECH - Speech Modeling for Facilitating Oral-Based Communication
Inria Nancy - Grand Est, LORIA - NLPKD - Department of Natural Language Processing & Knowledge Discovery
2 PAROLE - Analysis, perception and recognition of speech
INRIA Lorraine, LORIA - Laboratoire Lorrain de Recherche en Informatique et ses Applications
Abstract : We consider the problem of robust automatic speech recog-nition (ASR) in the context of the CHiME-3 Challenge. The proposed system combines three contributions. First, we propose a deep neural network (DNN) based multichannel speech enhancement technique, where the speech and noise spectra are estimated using a DNN based regressor and the spatial parameters are derived in an expectation-maximization (EM) like fashion. Second, a conditional restricted Boltz-mann machine (CRBM) model is trained using the obtained enhanced speech and used to generate simulated training and development datasets. The goal is to increase the similarity between simulated and real data, so as to increase the benefit of multicondition training. Finally, we make some changes to the ASR backend. Our system ranked 4th among 25 entries.
Type de document :
Communication dans un congrès
2015 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2015), Dec 2015, Arizona, United States
Liste complète des métadonnées

Littérature citée [39 références]  Voir  Masquer  Télécharger

https://hal.inria.fr/hal-01204553
Contributeur : Sunit Sivasankaran <>
Soumis le : jeudi 24 septembre 2015 - 11:54:07
Dernière modification le : mercredi 21 février 2018 - 07:50:03
Document(s) archivé(s) le : mardi 29 décembre 2015 - 09:43:47

Fichier

INRIA.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : hal-01204553, version 1

Citation

Sunit Sivasankaran, Aditya Arie Nugraha, Emmanuel Vincent, Juan Andrés Morales Cordovilla, Siddharth Dalmia, et al.. Robust ASR using neural network based speech enhancement and feature simulation. 2015 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2015), Dec 2015, Arizona, United States. 〈hal-01204553〉

Partager

Métriques

Consultations de la notice

736

Téléchargements de fichiers

1111