Integrating Data from the Web by Machine-Learning Tree-Pattern Queries - Inria - Institut national de recherche en sciences et technologies du numérique Accéder directement au contenu
Communication Dans Un Congrès Année : 2006

Integrating Data from the Web by Machine-Learning Tree-Pattern Queries

Résumé

Effienct and reliable integration of web data requires building programs called wrappers. Hand writting wrappers is tedious and error prone. Constant changes in the web, also implies that wrappers need to be constantly refactored. Machine learning has proven to be useful, but current techniques are either limited in expressivity, require non-intuitive user interaction or do not allow for n-ary extraction. We study using tree-patterns as an n-ary extraction language and propose an algorithm learning such queries. It calculates the most information-conservative tree-pattern which is a generalization of two input trees. A notable aspect is that the approach allows to learn queries containing both child and descendant relationships between nodes. More importantly, the proposed approach does not require any labeling other than the data which the user effectively wants to extract. The experiments reported show the effectiveness of the approach.
Fichier non déposé

Dates et versions

inria-00536547 , version 1 (16-11-2010)

Identifiants

  • HAL Id : inria-00536547 , version 1

Citer

Benjamin Habegger, Denis Debarbieux. Integrating Data from the Web by Machine-Learning Tree-Pattern Queries. 5th International Conference on Ontologies, Databases, and Applications of Semantics, 2006, Montpellier, France. pp.941-948. ⟨inria-00536547⟩
60 Consultations
0 Téléchargements

Partager

Gmail Facebook X LinkedIn More