Skip to Main content Skip to Navigation
Conference papers

DSS: A Scalable and Efficient Stratified Sampling Algorithm for Large-Scale Datasets

Abstract : Statistical analysis of aggregated records is widely used in various domains such as market research, sociological investigation and network analysis, etc. Stratified sampling (SS), which samples the population divided into distinct groups separately, is preferred in the practice for its high effectiveness and accuracy. In this paper, we propose a scalable and efficient algorithm named DSS, for SS to process large datasets. DSS executes all the sampling operations in parallel by calculating the exact subsample size for each partition according to the data distribution. We implement DSS on Spark, a big-data processing system, and we show through large-scale experiments that it can achieve lower data-transmission cost and higher efficiency than state-of-the-art methods with high sample representativeness.
Document type :
Conference papers
Complete list of metadata

Cited literature [21 references]  Display  Hide  Download

https://hal.inria.fr/hal-01648006
Contributor : Hal Ifip <>
Submitted on : Friday, November 24, 2017 - 4:49:17 PM
Last modification on : Tuesday, September 3, 2019 - 3:04:02 PM

File

432484_1_En_11_Chapter.pdf
Files produced by the author(s)

Licence


Distributed under a Creative Commons Attribution 4.0 International License

Identifiers

Citation

Minne Li, Dongsheng Li, Siqi Shen, Zhaoning Zhang, Xicheng Lu. DSS: A Scalable and Efficient Stratified Sampling Algorithm for Large-Scale Datasets. 13th IFIP International Conference on Network and Parallel Computing (NPC), Oct 2016, Xi'an, China. pp.133-146, ⟨10.1007/978-3-319-47099-3_11⟩. ⟨hal-01648006⟩

Share

Metrics

Record views

171

Files downloads

434