BERT and fastText Embeddings for Automatic Detection of Toxic Speech - Inria - Institut national de recherche en sciences et technologies du numérique Accéder directement au contenu
Communication Dans Un Congrès Année : 2020

BERT and fastText Embeddings for Automatic Detection of Toxic Speech

Résumé

With the expansion of Internet usage, catering to the dissemination of thoughts and expressions of an individual, there has been an immense increase in the spread of online hate speech. Social media, community forums, discussion platforms are few examples of common playground of online discussions where people are freely allowed to communicate. However, the freedom of speech may be misused by some people by arguing aggressively, offending others and spreading verbal violence. As there is no clear distinction between the terms offensive, abusive, hate and toxic speech, in this paper we consider the above mentioned terms as toxic speech. In many countries, online toxic speech is punishable by the law. Thus, it is important to automatically detect and remove toxic speech from online medias. Through this work, we propose automatic classification of toxic speech using embedding representations of words and deep-learning techniques. We perform binary and multi-class classification using a Twitter corpus and study two approaches: (a) a method which consists in extracting of word embeddings and then using a DNN classifier; (b) fine-tuning the pre-trained BERT model. We observed that BERT fine-tuning performed much better. Proposed methodology can be used for any other type of social media comments.
Fichier principal
Vignette du fichier
IEEE_SIIE2020_v4.pdf (356.73 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

hal-02448197 , version 1 (22-01-2020)
hal-02448197 , version 2 (01-04-2020)

Identifiants

  • HAL Id : hal-02448197 , version 2

Citer

Ashwin Geet d'Sa, Irina Illina, Dominique Fohr. BERT and fastText Embeddings for Automatic Detection of Toxic Speech. SIIE 2020 - Information Systems and Economic Intelligence; International Multi-Conference on:“Organization of Knowledge and Advanced Technologies”(OCTA), Feb 2020, Tunis, Tunisia. ⟨hal-02448197v2⟩
976 Consultations
6955 Téléchargements

Partager

Gmail Facebook X LinkedIn More