Discovering and Leveraging Content Similarity to Optimize Collective On-Demand Data Access to IaaS Cloud Storage - Archive ouverte HAL Access content directly
Conference Papers Year :

Discovering and Leveraging Content Similarity to Optimize Collective On-Demand Data Access to IaaS Cloud Storage

(1) , (2) , (2)
1
2
Bogdan Nicolae
Andrzej Kochut
  • Function : Author
  • PersonId : 965444
Alexei Karve
  • Function : Author
  • PersonId : 965445

Abstract

A critical feature of IaaS cloud computing is the ability to quickly disseminate the content of a shared dataset at large scale. In this context, a common pattern is collective on-demand read, i.e., accessing the same VM image or dataset from a large number of VM instances concurrently. There are various techniques that avoid I/O contention to the storage service where the dataset is located without relying on pre-broadcast. Most such techniques employ peer-to-peer collaborative behavior where the VM instances exchange information about the content that was accessed during runtime, such that it is possible to fetch the missing data pieces directly from each other rather than the storage system. However, such techniques are often limited within a group that performs a collective read. In light of high data redundancy on large IaaS data centers and multiple users that simultaneously run VM instance groups that perform collective reads, an important opportunity arises: enabling unrelated VM instances belonging to different groups to collaborate and exchange common data in order to further reduce the I/O pressure on the storage system. This paper deals with the challenges posed by such a solution, which prompt the need for novel techniques to efficiently detect and leverage common data pieces across groups. To this end, we introduce a low-overhead fingerprint based approach that we evaluate and demonstrate to be efficient in practice for a representative scenario on dozens of nodes and a variety of group configurations.
Fichier principal
Vignette du fichier
paper.pdf (184.8 Ko) Télécharger le fichier
Origin : Files produced by the author(s)
Loading...

Dates and versions

hal-01138684 , version 1 (02-04-2015)

Identifiers

  • HAL Id : hal-01138684 , version 1

Cite

Bogdan Nicolae, Andrzej Kochut, Alexei Karve. Discovering and Leveraging Content Similarity to Optimize Collective On-Demand Data Access to IaaS Cloud Storage. CCGrid'15: 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, May 2015, Shenzhen, China. ⟨hal-01138684⟩
139 View
186 Download

Share

Gmail Facebook Twitter LinkedIn More