Robust ASR  using neural network based speech enhancement and feature simulation

We consider the problem of robust automatic speech recognition (ASR) in the context of the CHiME-3 Challenge. The proposed system combines three contributions. First, we propose a deep neural network (DNN) based multichannel speech enhancement technique, where the speech and noise spectra are estimated using a DNN based regressor and the spatial parameters are derived in an expectation-maximization (EM) like fashion. Second, a conditional restricted Boltz-mann machine (CRBM) model is trained using the obtained enhanced speech and used to generate simulated training and development datasets. The goal is to increase the similarity between simulated and real data, so as to increase the benefit of multicondition training. Finally, we make some changes to the ASR backend. Our system ranked 4th among 25 entries

Mots clés

CHiME-3 CRBM feature simulation speech enhancement ASR

Domaines

Apprentissage [cs.LG] Son [cs.SD]

Fichier principal

INRIA.pdf (388.81 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Sunit Sivasankaran : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-01204553

Soumis le : jeudi 24 septembre 2015-11:54:07

Dernière modification le : jeudi 1 février 2024-10:03:33

Archivage à long terme le : mardi 29 décembre 2015-09:43:47

Dates et versions

hal-01204553 , version 1 (24-09-2015)

Identifiants

HAL Id : hal-01204553 , version 1

Citer

Sunit Sivasankaran, Aditya A Nugraha, Emmanuel Vincent, Juan Andrés Morales Cordovilla, Siddharth Dalmia, et al.. Robust ASR using neural network based speech enhancement and feature simulation. ASRU, Dec 2015, Arizona, United States. ⟨hal-01204553⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-RENNES1 CNRS INRIA IRISA GRID5000 UNIV-LORRAINE INRIA2 LORIA LORIA-NLPKD UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES SILECS UR1-MATH-NUM

532 Consultations

1379 Téléchargements