Rethinking Automatic Evaluation in Sentence Simplification

Thomas Scialom; Louis Martin; Jacopo Staiano; Eric Villemonte de La Clergerie; Benoît Sagot

Pré-Publication, Document De Travail Année : 2021

Rethinking Automatic Evaluation in Sentence Simplification

(1, 2) , (3, 4) , (1) , (4) , (4)

1
2
3
4

Thomas Scialom

Fonction : Auteur

reciTAL

Sorbonne Université

Louis Martin

Fonction : Auteur
PersonId : 171150
IdHAL : louis-martin
IdRef : 25975479X

Facebook AI Research [Paris]

Automatic Language Modelling and ANAlysis & Computational Humanities

Jacopo Staiano

Fonction : Auteur

reciTAL

Eric Villemonte de La Clergerie

Fonction : Auteur
PersonId : 1179
IdHAL : eric-villemonte-de-la-clergerie

Automatic Language Modelling and ANAlysis & Computational Humanities

Benoît Sagot

Fonction : Auteur
PersonId : 1461
IdHAL : bsagot
ORCID : 0000-0002-0107-8526
IdRef : 177454229

Automatic Language Modelling and ANAlysis & Computational Humanities

Résumé

Automatic evaluation remains an open research question in Natural Language Generation. In the context of Sentence Simplification, this is particularly challenging: the task requires by nature to replace complex words with simpler ones that shares the same meaning. This limits the effectiveness of n-gram based metrics like BLEU. Going hand in hand with the recent advances in NLG, new metrics have been proposed, such as BERTScore for Machine Translation. In summarization, the QuestEval metric proposes to automatically compare two texts by questioning them. In this paper, we first propose a simple modification of QuestEval allowing it to tackle Sentence Simplification. We then extensively evaluate the correlations w.r.t. human judgement for several metrics including the recent BERTScore and QuestEval, and show that the latter obtain state-of-the-art correlations, outperforming standard metrics like BLEU and SARI. More importantly, we also show that a large part of the correlations are actually spurious for all the metrics. To investigate this phenomenon further, we release a new corpus of evaluated simplifications, this time not generated by systems but instead, written by humans. This allows us to remove the spurious correlations and draw very different conclusions from the original ones, resulting in a better understanding of these metrics. In particular, we raise concerns about very low correlations for most of traditional metrics. Our results show that the only significant measure of the Meaning Preservation is our adaptation of QuestEval.

Domaines

Informatique et langage [cs.CL]

Benoît Sagot : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-03199901

Soumis le : vendredi 16 avril 2021-09:16:04

Dernière modification le : jeudi 1 février 2024-10:03:52

Dates et versions

hal-03199901 , version 1 (16-04-2021)

Identifiants

HAL Id : hal-03199901 , version 1
ARXIV : 2104.07560

Citer

Thomas Scialom, Louis Martin, Jacopo Staiano, Eric Villemonte de La Clergerie, Benoît Sagot. Rethinking Automatic Evaluation in Sentence Simplification. 2021. ⟨hal-03199901⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-RENNES1 INRIA IRISA INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES SORBONNE-UNIVERSITE UR1-MATH-NUM

60 Consultations

0 Téléchargements

Rethinking Automatic Evaluation in Sentence Simplification

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager