HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI Applications

Abstract : High performance computing will probably reach exascale in this decade. At this scale, mean time between failures is expected to be a few hours. Existing fault tolerant protocols for message passing applications will not be efficient anymore since they either require a global restart after a failure (checkpointing protocols) or result in huge memory occupation (message logging). Hybrid fault tolerant protocols overcome these limits by dividing applications processes into clusters and applying a different protocol within and between clusters. Combining coordinated checkpointing inside the clusters and message logging for the inter-cluster messages allows confining the consequences of a failure to a single cluster, while logging only a subset of the messages. However, in existing hybrid protocols, event logging is required for all application messages to ensure a correct execution after a failure. This can significantly impair failure free performance. In this paper, we propose HydEE, a hybrid rollback-recovery protocol for send-deterministic message passing applications, that provides failure containment without logging any event, and only a subset of the application messages. We prove that HydEE can handle multiple concurrent failures by relying on the send-deterministic execution model. Experimental evaluations of our implementation of HydEE in the MPICH2 library show that it introduces almost no overhead on failure free execution.
Complete list of metadatas

https://hal.inria.fr/hal-01121941
Contributor : Thomas Ropars <>
Submitted on : Monday, March 2, 2015 - 9:26:20 PM
Last modification on : Thursday, August 1, 2019 - 2:12:06 PM
Long-term archiving on : Tuesday, June 2, 2015 - 9:55:51 AM

File

ipdps2012.pdf
Files produced by the author(s)

Identifiers

Citation

Amina Guermouche, Thomas Ropars, Marc Snir, Franck Cappello. HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI Applications. , 2012, Shanghai, China. ⟨10.1109/IPDPS.2012.111⟩. ⟨hal-01121941⟩

Share

Metrics

Record views

253

Files downloads

604