Skip to Main content Skip to Navigation
New interface
Conference papers

HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI Applications

Abstract : High performance computing will probably reach exascale in this decade. At this scale, mean time between failures is expected to be a few hours. Existing fault tolerant protocols for message passing applications will not be efficient anymore since they either require a global restart after a failure (checkpointing protocols) or result in huge memory occupation (message logging). Hybrid fault tolerant protocols overcome these limits by dividing applications processes into clusters and applying a different protocol within and between clusters. Combining coordinated checkpointing inside the clusters and message logging for the inter-cluster messages allows confining the consequences of a failure to a single cluster, while logging only a subset of the messages. However, in existing hybrid protocols, event logging is required for all application messages to ensure a correct execution after a failure. This can significantly impair failure free performance. In this paper, we propose HydEE, a hybrid rollback-recovery protocol for send-deterministic message passing applications, that provides failure containment without logging any event, and only a subset of the application messages. We prove that HydEE can handle multiple concurrent failures by relying on the send-deterministic execution model. Experimental evaluations of our implementation of HydEE in the MPICH2 library show that it introduces almost no overhead on failure free execution.
Complete list of metadata
Contributor : Thomas Ropars Connect in order to contact the contributor
Submitted on : Monday, March 2, 2015 - 9:26:20 PM
Last modification on : Friday, October 7, 2022 - 3:48:29 AM
Long-term archiving on: : Tuesday, June 2, 2015 - 9:55:51 AM


Files produced by the author(s)



Amina Guermouche, Thomas Ropars, Marc Snir, Franck Cappello. HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI Applications. {26th IEEE International Parallel & Distributed Processing Symposium (IPDPS2012), 2012, Shanghai, China. ⟨10.1109/IPDPS.2012.111⟩. ⟨hal-01121941⟩



Record views


Files downloads