Improving Short Text Classification Through Global Augmentation Methods

Vukosi Marivate; Tshephisho Sefara

doi:10.1007/978-3-030-57321-8_21

Communication Dans Un Congrès Année : 2020

Improving Short Text Classification Through Global Augmentation Methods

(1, 2) , (3)

1
2
3

Vukosi Marivate

Fonction : Auteur
PersonId : 1098364

University of Pretoria [South Africa]

Council for Scientific and Industrial Research [South Africa]

Tshephisho Sefara

Fonction : Auteur
PersonId : 1115839

Council for Scientific and Industrial Research [Pretoria]

Résumé

We study the effect of different approaches to text augmentation. To do this we use three datasets that include social media and formal text in the form of news articles. Our goal is to provide insights for practitioners and researchers on making choices for augmentation for classification use cases. We observe that Word2Vec-based augmentation is a viable option when one does not have access to a formal synonym model (like WordNet-based augmentation). The use of mixup further improves performance of all text based augmentations and reduces the effects of overfitting on a tested deep learning model. Round-trip translation with a translation service proves to be harder to use due to cost and as such is less accessible for both normal and low resource use-cases.

Mots clés

Natural Language Processing Data augmentation Deep Neural Networks Text classification

Domaines

Informatique [cs] Sciences de l'information et de la communication

Fichier principal

497121_1_En_21_Chapter.pdf (532.55 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Hal Ifip : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-03414750

Soumis le : jeudi 4 novembre 2021-15:58:34

Dernière modification le : vendredi 5 novembre 2021-03:57:59

Archivage à long terme le : samedi 5 février 2022-19:10:49

Dates et versions

hal-03414750 , version 1 (04-11-2021)

Licence

Paternité

Identifiants

HAL Id : hal-03414750 , version 1
DOI : 10.1007/978-3-030-57321-8_21

Citer

Vukosi Marivate, Tshephisho Sefara. Improving Short Text Classification Through Global Augmentation Methods. 4th International Cross-Domain Conference for Machine Learning and Knowledge Extraction (CD-MAKE), Aug 2020, Dublin, Ireland. pp.385-399, ⟨10.1007/978-3-030-57321-8_21⟩. ⟨hal-03414750⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

IFIP-LNCS IFIP IFIP-TC IFIP-TC5 IFIP-WG IFIP-TC12 IFIP-TC8 IFIP-WG8-4 IFIP-WG8-9 IFIP-CD-MAKE IFIP-WG12-9 IFIP-LNCS-12279

73 Consultations

11 Téléchargements

Improving Short Text Classification Through Global Augmentation Methods

Résumé

Mots clés

Domaines

Dates et versions

Licence

Identifiants

Citer

Exporter

Collections

Altmetric

Partager