Benchmarking the Memory Hierarchy of Modern GPUs

Abstract : Memory access efficiency is a key factor for fully exploiting the computational power of Graphics Processing Units (GPUs). However, many details of the GPU memory hierarchy are not released by the vendors. We propose a novel fine-grained benchmarking approach and apply it on two popular GPUs, namely Fermi and Kepler, to expose the previously unknown characteristics of their memory hierarchies. Specifically, we investigate the structures of different cache systems, such as data cache, texture cache, and the translation lookaside buffer (TLB). We also investigate the impact of bank conflict on shared memory access latency. Our benchmarking results offer a better understanding on the mysterious GPU memory hierarchy, which can help in the software optimization and the modelling of GPU architectures. Our source code and experimental results are publicly available.
Type de document :
Communication dans un congrès
Ching-Hsien Hsu; Xuanhua Shi; Valentina Salapura. 11th IFIP International Conference on Network and Parallel Computing (NPC), Sep 2014, Ilan, Taiwan. Springer, Lecture Notes in Computer Science, LNCS-8707, pp.144-156, 2014, Network and Parallel Computing. 〈10.1007/978-3-662-44917-2_13〉
Liste complète des métadonnées

Littérature citée [13 références]  Voir  Masquer  Télécharger

https://hal.inria.fr/hal-01403075
Contributeur : Hal Ifip <>
Soumis le : vendredi 25 novembre 2016 - 14:28:09
Dernière modification le : vendredi 1 décembre 2017 - 01:10:04

Fichier

978-3-662-44917-2_13_Chapter.p...
Fichiers produits par l'(les) auteur(s)

Licence


Distributed under a Creative Commons Paternité 4.0 International License

Identifiants

Citation

Xinxin Mei, Kaiyong Zhao, Chengjian Liu, Xiaowen Chu. Benchmarking the Memory Hierarchy of Modern GPUs. Ching-Hsien Hsu; Xuanhua Shi; Valentina Salapura. 11th IFIP International Conference on Network and Parallel Computing (NPC), Sep 2014, Ilan, Taiwan. Springer, Lecture Notes in Computer Science, LNCS-8707, pp.144-156, 2014, Network and Parallel Computing. 〈10.1007/978-3-662-44917-2_13〉. 〈hal-01403075〉

Partager

Métriques

Consultations de la notice

40

Téléchargements de fichiers

34