Active Optimistic Message Logging for Reliable Execution of MPI Applications - Inria - Institut national de recherche en sciences et technologies du numérique Access content directly
Conference Papers Year : 2009

Active Optimistic Message Logging for Reliable Execution of MPI Applications

Abstract

To execute MPI applications reliably, fault tolerance mechanisms are needed. Message logging is a well known solution to provide fault tolerance for MPI applications. It as been proved that it can tolerate higher failure rate than coordinated checkpointing. However pessimistic and causal message logging can induce high overhead on failure free execution. In this paper, we present O2P, a new optimistic message logging protocol, based on active optimistic message logging. Contrary to existing optimistic message logging protocols that saves dependency information on reliable storage periodically, O2P logs dependency information as soon as possible to reduce the amount of data piggybacked on application messages. Thus it reduces the overhead of the protocol on failure free execution, making it more scalable and simplifying recovery. O2P is implemented as a module of the Open MPI library. Experiments show that active message logging is promising to improve scalability and performance of optimistic message logging.
No file

Dates and versions

inria-00424002 , version 1 (13-10-2009)

Identifiers

  • HAL Id : inria-00424002 , version 1

Cite

Thomas Ropars, Christine Morin. Active Optimistic Message Logging for Reliable Execution of MPI Applications. 15th International Euro-Par Conference, Aug 2009, Delft, Netherlands. ⟨inria-00424002⟩
231 View
0 Download

Share

Gmail Facebook X LinkedIn More