Ungoliant: An Optimized Pipeline for the Generation of a Very Large-Scale Multilingual Web Corpus

Since the introduction of large language models in Natural Language Processing, large raw corpora have played a crucial role in Computational Linguistics. However, most of these large raw corpora are either available only for English or not available to the general public due to copyright issues. Nevertheless, there are some examples of freely available multilingual corpora for training Deep Learning NLP models, such as the OSCAR and Paracrawl corpora. However, they have quality issues, especially for low-resource languages. Moreover, recreating or updating these corpora is very complex. In this work, we try to reproduce and improve the goclassy pipeline used to create the OSCAR corpus. We propose a new pipeline that is faster, modular, parameterizable, and well documented. We use it to create a corpus similar to OSCAR but larger and based on recent data.Also, unlike OSCAR, the metadata information is at the document level. We release our pipeline under an open source license and publish the corpus under a research-only license.

Domaines

Informatique et langage [cs.CL]

Fichier principal

Ungoliant___CMLC_9-2.pdf (236.68 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Pedro Ortiz Suarez : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-03301590

Soumis le : mardi 27 juillet 2021-15:27:40

Dernière modification le : jeudi 1 février 2024-10:06:29

Archivage à long terme le : jeudi 28 octobre 2021-18:30:01

Dates et versions

hal-03301590 , version 1 (27-07-2021)

Licence

Paternité

Identifiants

HAL Id : hal-03301590 , version 1
DOI : 10.14618/ids-pub-10468

Citer

Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, Benoît Sagot. Ungoliant: An Optimized Pipeline for the Generation of a Very Large-Scale Multilingual Web Corpus. CMLC 2021 - 9th Workshop on Challenges in the Management of Large Corpora, Jul 2021, Limerick / Virtual, Ireland. ⟨10.14618/ids-pub-10468⟩. ⟨hal-03301590⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-RENNES1 INRIA IRISA INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES SORBONNE-UNIVERSITE ANR PRAIRIE-IA UR1-MATH-NUM

229 Consultations

310 Téléchargements