The Failure Trace Archive: Enabling Comparative Analysis of Failures in Diverse Distributed Systems

Derrick Kondo 1 Bahman Javadi 1 Alexandru Iosup 2 Dick Epema 2
1 MESCAL - Middleware efficiently scalable
Inria Grenoble - Rhône-Alpes, LIG - Laboratoire d'Informatique de Grenoble
Abstract : With the increasing functionality and complexity of distributed systems, resource failures are inevitable. While numerous models and algorithms for dealing with failures exist, the lack of public trace data sets and tools have prevented meaningful comparisons. To facilitate the design, validation, and comparison of fault-tolerant models and algorithms, we have created the Failure Trace Archive (FTA) as an online public repository of availability traces taken from diverse parallel and distributed systems. Our main contributions in this study are the following. First, we describe the design of the archive, in particular the rationale of the standard FTA format, and the design of a toolbox that facilitates automated analysis of trace data sets. Second, applying the toolbox, we present a uniform comparative analysis with statistics and models of failures in nine distributed systems. Third, we show how different interpretations of these data sets can result in different conclusions. This emphasizes the critical need for the public availability of trace data and methods for their analysis.
Type de document :
Rapport
[Research Report] 2009
Liste complète des métadonnées

Littérature citée [22 références]  Voir  Masquer  Télécharger

https://hal.inria.fr/inria-00433523
Contributeur : Derrick Kondo <>
Soumis le : jeudi 19 novembre 2009 - 16:50:32
Dernière modification le : mercredi 11 avril 2018 - 01:52:07
Document(s) archivé(s) le : mardi 16 octobre 2012 - 14:30:32

Fichier

ccgrid10.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : inria-00433523, version 1

Collections

Citation

Derrick Kondo, Bahman Javadi, Alexandru Iosup, Dick Epema. The Failure Trace Archive: Enabling Comparative Analysis of Failures in Diverse Distributed Systems. [Research Report] 2009. 〈inria-00433523〉

Partager

Métriques

Consultations de la notice

380

Téléchargements de fichiers

623