Preparing the Dictionnaire Universel for Automatic Enrichment

Pedro Javier Ortiz Suárez; Laurent Romary; Benoît Sagot

Communication Dans Un Congrès Année : 2019

Preparing the Dictionnaire Universel for Automatic Enrichment

(1, 2) , (1) , (1)

1
2

Pedro Javier Ortiz Suárez

Fonction : Auteur
PersonId : 178412
IdHAL : pedro-ortiz-suarez
ORCID : 0000-0003-0343-8852
IdRef : 264210743

Automatic Language Modelling and ANAlysis & Computational Humanities

Sorbonne Université

Laurent Romary

Fonction : Auteur
PersonId : 307
IdHAL : laurentromary
ORCID : 0000-0002-0756-0508
IdRef : 060702494

Automatic Language Modelling and ANAlysis & Computational Humanities

Benoît Sagot

Fonction : Auteur
PersonId : 1461
IdHAL : bsagot
ORCID : 0000-0002-0107-8526
IdRef : 177454229

Automatic Language Modelling and ANAlysis & Computational Humanities

Résumé

The Dictionnaire Universel (DU) is an encyclopaedic dictionary originally written by Antoine Furetière around 1676-78, later revised and improved by the Protestant jurist Henri Basnage de Beauval who expanded, corrected and included terms of arts, crafts and sciences, into the Dictionnaire. The aim of the BASNUM project is to digitize the DU in its second edition rewritten by Basnage de Beauval, to analyse it with computational methods in order to better assess the importance of this work for the evolution of sciences and mentalities in the 18th century, and to contribute to the contemporary movement for creating innovative and data-driven computational methods for text digitization, encoding and analysis. Based on the experience acquired within the research group, an enrichment workflow based upon a series of Natural Language Processing processes is being set up to be applied to Basnage's work. This includes, among others, automatic identification of the dictionary structure (macro-, meso- and microstructure), named-entity recognition (in particular persons and locations), classification of dictionary entries, detection and study of polysemy markers, tracking and classification of quotation use (bibliographic references), scoring semantic similarity between the DU and other dictionaries. The main challenges being the lack of available annotated data in order to train machine learning models, decreased accuracy when using modern pre-trained models due to the differences between present-day and 18th century French, and even unreliable or low quality OCRisation. The paper describes methods that are useful to tackle these issues in order to prepare the the DU for automatic enrichment going beyond what current available tools like Grobid-dictionaries can do, thanks to the advent of deep learning NLP models. The paper also describes how these methods could be applied to other dictionaries or even other types of ancient texts.

Domaines

Traitement du texte et du document Informatique et langage [cs.CL] Linguistique Héritage culturel et muséologie

Fichier principal

ICHLL_10_Slides.pdf (3.68 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Benoît Sagot : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-02131598

Soumis le : vendredi 18 octobre 2019-19:36:19

Dernière modification le : jeudi 1 février 2024-10:03:37

Archivage à long terme le : dimanche 19 janvier 2020-14:31:14

Dates et versions

hal-02131598 , version 1 (18-10-2019)

Identifiants

HAL Id : hal-02131598 , version 1

Citer

Pedro Javier Ortiz Suárez, Laurent Romary, Benoît Sagot. Preparing the Dictionnaire Universel for Automatic Enrichment. 10th International Conference on Historical Lexicography and Lexicology (ICHLL), Jun 2019, Leeuwarden, Netherlands. ⟨hal-02131598⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-RENNES1 INRIA IRISA INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES SORBONNE-UNIVERSITE ANR UR1-MATH-NUM

200 Consultations

79 Téléchargements

Preparing the Dictionnaire Universel for Automatic Enrichment

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager