Learnable Nonlinear Compression for Robust Speaker Verification - Inria - Institut national de recherche en sciences et technologies du numérique Accéder directement au contenu
Communication Dans Un Congrès Année : 2022

Learnable Nonlinear Compression for Robust Speaker Verification

Résumé

In this study, we focus on nonlinear compression methods in spectral features for speaker verification based on deep neural network. We consider different kinds of channel-dependent (CD) nonlinear compression methods optimized in a data-driven manner. Our methods are based on power nonlinearities and dynamic range compression (DRC). We also propose multi-regime (MR) design on the nonlinearities, at improving robustness. Results on VoxCeleb1 and Vox-Movies data demonstrate improvements brought by proposed compression methods over both the commonly-used logarithm and their static counterparts, especially for ones based on power function. While CD generalization improves performance on VoxCeleb1, MR provides more robustness on VoxMovies, with a maximum relative equal error rate reduction of 21.6%.
Fichier principal
Vignette du fichier
LearnableNonlinear_ICASSP2022.pdf (2 Mo) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)

Dates et versions

hal-03616852 , version 1 (23-03-2022)

Identifiants

Citer

Xuechen Liu, Md Sahidullah, Tomi Kinnunen. Learnable Nonlinear Compression for Robust Speaker Verification. ICASSP 2022 - IEEE International Conference on Acoustics, Speech and Signal Processing, May 2022, Singapore, Singapore. ⟨10.1109/ICASSP43922.2022.9747185⟩. ⟨hal-03616852⟩
49 Consultations
64 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More