Large-scale experiment for topology-aware resource management - Inria - Institut national de recherche en sciences et technologies du numérique Accéder directement au contenu
Communication Dans Un Congrès Année : 2018

Large-scale experiment for topology-aware resource management

Résumé

A Resource and Job Management System (RJMS) is a crucial system software part of the HPC stack. It is responsible for efficiently delivering computing power to applications in supercomputing environments and its main intelligence relies on resource selection techniques to find the most adapted resources to schedule the users' jobs. In [8], we introduced a new topology-aware resource selection algorithm to determine the best choice among the available nodes of the platform based on their position in the network and on application behaviour (expressed as a communication matrix). We did integrate this algorithm as a plugin in Slurm and validated it with several optimization schemes by making comparisons with the default Slurm algorithm. This paper presents further experiments with regard to this selection process.
Fichier principal
Vignette du fichier
article_89.pdf (285.65 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

hal-01667350 , version 1 (19-12-2017)

Identifiants

  • HAL Id : hal-01667350 , version 1

Citer

Yiannis Georgiou, Guillaume Mercier, Adèle Villiermet. Large-scale experiment for topology-aware resource management. Open workshop on data locality, Aug 2017, Santiago de Compostella, Spain. ⟨hal-01667350⟩
190 Consultations
203 Téléchargements

Partager

Gmail Facebook X LinkedIn More