SanskritTagger : a stochastic lexical and pos tagger for Sanskrit

Abstract : SanskritTagger is a stochastic tagger for unpreprocessed Sanskrit text. The tagger tokenises text with a Markov model and performs part-of-speech tagging with a Hidden Markov model. Parameters for these processes are estimated from a manually annotated corpus of currently about 1.500.000 words. The article sketches the tagging process, reports the results of tagging a few short passages of Sanskrit text and describes further improvements of the program. The article describes design and function of SanskritTagger, a tokeniser and part-of-speech (POS) tagger, which analyses ”natural”, i.e. unannotated Sanskrit text by repeated application of stochastic models. This tagger has been developped during the last few years as part of a larger project for digitalisation of Sanskrit texts (cmp. (Hellwig, 2002)) and is still in the state of steady improvement. The article is organised as follows: Section 1 gives a short overview about linguistic problems found in Sanskrit texts which influenced the design of the tagger. Section 2 describes the actual implementation of the tagger. In section 3, the performance of the tagger is evaluated on short passages of text from different thematic areas. In addition, this section describes possible improvements in future versions.
Type de document :
Communication dans un congrès
Gérard Huet and Amba Kulkarni. First International Sanskrit Computational Linguistics Symposium, Oct 2007, Rocquencourt, France. 2007, http://hal.inria.fr/SANSKRIT/fr/
Liste complète des métadonnées

Littérature citée [6 références]  Voir  Masquer  Télécharger

https://hal.inria.fr/inria-00203467
Contributeur : Brigitte Briot <>
Soumis le : jeudi 10 janvier 2008 - 11:24:33
Dernière modification le : jeudi 28 février 2008 - 15:52:42
Document(s) archivé(s) le : mardi 13 avril 2010 - 16:54:28

Fichier

Hellwig.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : inria-00203467, version 1

Collections

Citation

Oliver Hellwig. SanskritTagger : a stochastic lexical and pos tagger for Sanskrit. Gérard Huet and Amba Kulkarni. First International Sanskrit Computational Linguistics Symposium, Oct 2007, Rocquencourt, France. 2007, http://hal.inria.fr/SANSKRIT/fr/. 〈inria-00203467〉

Partager

Métriques

Consultations de la notice

202

Téléchargements de fichiers

132