Approximated Summarization of Data Provenance

Ainy Eleanor 1 Pierre Bourhis 2 Susan Davidson 3 Daniel Deutch 1 Tova Milo 1
2 LINKS - Linking Dynamic Data
Inria Lille - Nord Europe, CRIStAL - Centre de Recherche en Informatique, Signal et Automatique de Lille (CRIStAL) - UMR 9189
Abstract : Many modern applications involve collecting large amounts of data from multiple sources, and then aggregating and manipulating it in intricate ways. The complexity of such applications, combined with the size of the collected data, makes it difficult to understand how the resulting information was derived. Data provenance has proven helpful in this respect, however, maintaining and presenting the full and exact provenance information may be infeasible due to its size and complexity. We therefore introduce the notion of approximated summarized provenance, which provides a compact representation of the provenance at the possible cost of information loss. Based on this notion, we present a novel provenance summarization algorithm which, based on the semantics of the underlying data and the intended use of provenance, outputs a summary of the input provenance. Experiments measure the conciseness and accuracy of the resulting provenance summaries, and improvement in provenance usage time.
Type de document :
Communication dans un congrès
CIKM, Oct 2015, Melbourn, Australia. 2015, 〈http://www.cikm-2015.org/〉
Liste complète des métadonnées

https://hal.inria.fr/hal-01211286
Contributeur : Inria Links <>
Soumis le : dimanche 4 octobre 2015 - 17:52:59
Dernière modification le : jeudi 11 janvier 2018 - 06:27:32

Identifiants

  • HAL Id : hal-01211286, version 1

Citation

Ainy Eleanor, Pierre Bourhis, Susan Davidson, Daniel Deutch, Tova Milo. Approximated Summarization of Data Provenance. CIKM, Oct 2015, Melbourn, Australia. 2015, 〈http://www.cikm-2015.org/〉. 〈hal-01211286〉

Partager

Métriques

Consultations de la notice

182