Skip to Main content Skip to Navigation
New interface
Journal articles

TomusBlobs: Scalable Data-intensive Processing on Azure Clouds

Alexandru Costan 1 Radu Tudoran 1 Gabriel Antoniu 1 Goetz Brasche 2 
1 KerData - Scalable Storage for Clouds and Beyond
Inria Rennes – Bretagne Atlantique , IRISA-D1 - SYSTÈMES LARGE ÉCHELLE
Abstract : The emergence of cloud computing has brought the opportunity to use large-scale compute infrastructures for a broader and broader spectrum of applications and users. As the cloud paradigm gets attractive for the "elasticity'' in resource usage and associated costs (the users only pay for resources actually used), cloud applications still suffer from the high latencies and low performance of cloud storage services. As Big Data analysis on clouds becomes more and more relevant in many application areas, enabling high-throughput massive data processing on cloud data becomes a critical issue, as it impacts the overall application performance. In this paper we address this challenge at the level of cloud storage. We introduce a concurrency-optimized data storage system (called TomusBlobs) which federates the virtual disks associated to the Virtual Machines running the application code on the cloud. We demonstrate the performance benefits of our solution for efficient data-intensive processing by building an optimized prototype MapReduce framework for Microsoft's Azure cloud platform based on TomusBlobs. Finally, we specifically address the limitations of state-of-the-art MapReduce frameworks for reduce-intensive workloads, by proposing MapIterativeReduce as an extension of the MapReduce model. We validate the above contributions through large-scale experiments with synthetic benchmarks and with real-world applications on the Azure commercial cloud, using resources distributed across multiple data centers: they demonstrate that our solutions bring substantial benefits to data intensive applications compared to approaches relying on state-of-the-art cloud object storage.
Complete list of metadata

Cited literature [22 references]  Display  Hide  Download
Contributor : Gabriel Antoniu Connect in order to contact the contributor
Submitted on : Wednesday, August 31, 2016 - 3:05:14 PM
Last modification on : Friday, November 18, 2022 - 9:27:02 AM
Long-term archiving on: : Friday, December 2, 2016 - 3:21:57 AM


Files produced by the author(s)



Alexandru Costan, Radu Tudoran, Gabriel Antoniu, Goetz Brasche. TomusBlobs: Scalable Data-intensive Processing on Azure Clouds. Concurrency and Computation: Practice and Experience, 2013, ⟨10.1002/cpe.3034⟩. ⟨hal-00767034⟩



Record views


Files downloads