Building a Large Syntactically-Annotated Corpus of Vietnamese - Inria - Institut national de recherche en sciences et technologies du numérique Accéder directement au contenu
Communication Dans Un Congrès Année : 2009

Building a Large Syntactically-Annotated Corpus of Vietnamese

Résumé

Treebank is an important resource for both research and application of natural language processing. For Vietnamese, we still lack such kind of corpora. This paper presents up-to-date results of a project for Vietnamese treebank construction. Since Vietnamese is an isolating language and has no word delimiter, there are many ambiguities in sentence analysis. We systematically applied a lot of linguistic techniques to handle such ambiguities. Annotators are supported by automatic labeling tools and a tree-editor tool. Raw texts are extracted from Tuoi Tre (Youth), an online Vietnamese daily newspaper. The current annotation agreement is around 90 percent.
Fichier principal
Vignette du fichier
LAW09-VTB-Thai.pdf (235.9 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)

Dates et versions

inria-00421103 , version 1 (12-11-2009)
inria-00421103 , version 2 (15-12-2009)

Identifiants

  • HAL Id : inria-00421103 , version 1

Citer

Hong Phuong Le, Thi Minh Huyen Nguyen, Phuong Thai Nguyen, Xuan Luong Vu, van Hiep Nguyen. Building a Large Syntactically-Annotated Corpus of Vietnamese. The Third Linguistic Annotation Workshop (The LAW III), Aug 2009, Singapour, Singapore. 6p. ⟨inria-00421103v1⟩
700 Consultations
1346 Téléchargements

Partager

Gmail Facebook X LinkedIn More