Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?

Brij Mohan Lal Srivastava; Aurélien Bellet; Marc Tommasi; Emmanuel Vincent

Communication Dans Un Congrès Année : 2019

Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?

(1) , (1) , (1) , (2)

1
2

Brij Mohan Lal Srivastava

Fonction : Auteur

Machine Learning in Information Networks

Aurélien Bellet

Fonction : Auteur
PersonId : 9877
IdHAL : aurelien-bellet
ORCID : 0000-0003-3440-1251
IdRef : 17653136X

Machine Learning in Information Networks

Marc Tommasi

Fonction : Auteur
PersonId : 399
IdHAL : marc-tommasi
ORCID : 0000-0003-2838-4408
IdRef : 121846385

Machine Learning in Information Networks

Emmanuel Vincent

Fonction : Auteur
PersonId : 1256
IdHAL : emmanuelv
ORCID : 0000-0002-0183-7289
IdRef : 089360176

Speech Modeling for Facilitating Oral-Based Communication

Résumé

Automatic speech recognition (ASR) is a key technology in many services and applications. This typically requires user devices to send their speech data to the cloud for ASR decoding. As the speech signal carries a lot of information about the speaker, this raises serious privacy concerns. As a solution, an encoder may reside on each user device which performs local computations to anonymize the representation. In this paper, we focus on the protection of speaker identity and study the extent to which users can be recognized based on the encoded representation of their speech as obtained by a deep encoder-decoder architecture trained for ASR. Through speaker identification and verification experiments on the Librispeech corpus with open and closed sets of speakers, we show that the representations obtained from a standard architecture still carry a lot of information about speaker identity. We then propose to use adversarial training to learn representations that perform well in ASR while hiding speaker identity. Our results demonstrate that adversarial training dramatically reduces the closed-set classification accuracy, but this does not translate into increased open-set verification error hence into increased protection of the speaker identity in practice. We suggest several possible reasons behind this negative result.

Mots clés

Speaker recognition Adversarial training End-to-end system Privacy Speech recognition

Domaines

Apprentissage [cs.LG] Machine Learning [stat.ML] Informatique et langage [cs.CL]

Fichier principal

srivastava_IS19.pdf (444.55 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Aurélien Bellet : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-02166434

Soumis le : mercredi 3 juillet 2019-09:59:13

Dernière modification le : jeudi 1 février 2024-10:03:27

Dates et versions

hal-02166434 , version 1 (03-07-2019)

Identifiants

HAL Id : hal-02166434 , version 1

Citer

Brij Mohan Lal Srivastava, Aurélien Bellet, Marc Tommasi, Emmanuel Vincent. Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?. INTERSPEECH 2019 - 20th Annual Conference of the International Speech Communication Association, Sep 2019, Graz, Austria. ⟨hal-02166434⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-RENNES1 CNRS INRIA IRISA GRID5000 CRISTAL UNIV-LORRAINE INRIA2 CRISTAL-MAGNET LORIA LORIA-NLPKD UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES UNIV-LILLE SILECS HYAIAI ANR UR1-MATH-NUM

295 Consultations

475 Téléchargements

Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager