Skip to Main content Skip to Navigation
Conference papers

Enhancing Content-And-Structure Information Retrieval using a Native XML Database

Abstract : Three approaches to content-and-structure XML retrieval are analysed in this paper: first by using Zettair, a full-text information retrieval system; second by using eXist, a native XML database, and third by using a hybrid XML retrieval system that uses eXist to produce the final answers from likely relevant articles retrieved by Zettair. INEX 2003 content-and-structure topics can be classified in two categories: the first retrieving full articles as final answers, and the second retrieving more specific elements within articles as final answers. We show that for both topic categories our initial hybrid system improves the retrieval effectiveness of a native XML database. For ranking the final answer elements, we propose and evaluate a novel retrieval model that utilises the structural relationships between the answer elements of a native XML database and retrieves Coherent Retrieval Elements. The final results of our experiments show that when the XML retrieval task focusses on highly relevant elements our hybrid XML retrieval system with the Coherent Retrieval Elements module is 1.8 times more effective than Zettair and 3 times more effective than eXist, and yields an effective content-and-structure XML retrieval.
Document type :
Conference papers
Complete list of metadata

Cited literature [12 references]  Display  Hide  Download

https://hal.inria.fr/inria-00000185
Contributor : Anne-Marie Vercoustre <>
Submitted on : Tuesday, August 2, 2005 - 5:00:36 PM
Last modification on : Thursday, February 11, 2021 - 2:50:06 PM
Long-term archiving on: : Thursday, April 1, 2010 - 10:10:53 PM

Identifiers

Collections

Citation

Jovan Pehcevski, James Thom, Anne-Marie Vercoustre. Enhancing Content-And-Structure Information Retrieval using a Native XML Database. The First Twente Data Management Workshop - TDM'04, Jun 2004, University of Twente, Enschede, The Netherlands. ⟨inria-00000185⟩

Share

Metrics

Record views

142

Files downloads

243