Duration modelling and evaluation for Arabic statistical parametric speech synthesis - Inria - Institut national de recherche en sciences et technologies du numérique Accéder directement au contenu
Article Dans Une Revue Multimedia Tools and Applications Année : 2020

Duration modelling and evaluation for Arabic statistical parametric speech synthesis

Résumé

Sound duration is responsible for rhythm and speech rate. Furthermore , in some languages phoneme length is an important phonetic and prosodic factor. For example, in Arabic, gemination and vowel quantity are two important characteristics of the language. Therefore, accurate duration modelling is crucial for Arabic TTS systems. This paper is interested in improving the modelling of phone duration for Arabic statistical parametric speech synthesis using DNN-based models. In fact, since a few years, DNN have been frequently used for parametric speech synthesis, instead of HMM. Therefore, several variants of DNN-based duration models for Arabic are investigated. The novelty consists in training a specific DNN model for each class of sounds, i.e. short vowels, long vowels, simple consonants and geminated consonants. The main idea behind this choice is the improvement that we already achieved in the quality of Arabic parametric speech synthesis by the introduction of two specific features of Arabic, i.e. gemination and vowel quantity into the standard HTS feature set. Both objective and subjective evaluations show that using a specific model for each class of sounds leads to a more accurate modelling of the phone duration in Arabic parametric speech synthesis, outperforming the state-of-the-art duration modelling systems.
Fichier principal
Vignette du fichier
Duration_Modelling_article_Sept_2020_Rev.pdf (671.87 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

hal-03007287 , version 1 (16-11-2020)

Identifiants

Citer

Imene Zangar, Zied Mnasri, Vincent Colotte, Denis Jouvet. Duration modelling and evaluation for Arabic statistical parametric speech synthesis. Multimedia Tools and Applications, 2020, ⟨10.1007/s11042-020-09901-7⟩. ⟨hal-03007287⟩
58 Consultations
225 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More