SPARQLGX: Efficient Distributed Evaluation of SPARQL with Apache Spark

Damien Graux; Louis Jachiet; Pierre Genevès; Nabil Layaïda

doi:10.1007/978-3-319-46547-0_9

Communication Dans Un Congrès Année : 2016

SPARQLGX: Efficient Distributed Evaluation of SPARQL with Apache Spark

(1) , (1) , (1) , (1)

Damien Graux

Fonction : Auteur
PersonId : 776616
ORCID : 0000-0003-3392-3162

Types and Reasoning for the Web

Louis Jachiet

Fonction : Auteur
PersonId : 179861
IdHAL : louis-jachiet

Types and Reasoning for the Web

Pierre Genevès

Fonction : Auteur correspondant
PersonId : 9676
IdHAL : pierre-geneves
ORCID : 0000-0001-7676-2755
IdRef : 117936324

Connectez-vous pour contacter l'auteur

Types and Reasoning for the Web

Nabil Layaïda

Fonction : Auteur
PersonId : 21665
IdHAL : nabil-layaida
ORCID : 0000-0001-8472-9365
IdRef : 15031504X

Types and Reasoning for the Web

Résumé

sparql is the w3c standard query language for querying data expressed in the Resource Description Framework (rdf). The increasing amounts of rdf data available raise a major need and research interest in building efficient and scalable distributed sparql query eval-uators. In this context, we propose sparqlgx: our implementation of a distributed rdf datastore based on Apache Spark. sparqlgx is designed to leverage existing Hadoop infrastructures for evaluating sparql queries. sparqlgx relies on a translation of sparql queries into exe-cutable Spark code that adopts evaluation strategies according to (1) the storage method used and (2) statistics on data. We show that spar-qlgx makes it possible to evaluate sparql queries on billions of triples distributed across multiple nodes, while providing attractive performance figures. We report on experiments which show how sparqlgx compares to related state-of-the-art implementations and we show that our approach scales better than these systems in terms of supported dataset size. With its simple design, sparqlgx represents an interesting alternative in several scenarios.

Domaines

Web

Fichier principal

sparqlgx.pdf (222.02 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Tyrex Equipe : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-01344915

Soumis le : mardi 12 juillet 2016-18:57:25

Dernière modification le : jeudi 4 avril 2024-21:06:50

Dates et versions

hal-01344915 , version 1 (12-07-2016)

Identifiants

HAL Id : hal-01344915 , version 1
DOI : 10.1007/978-3-319-46547-0_9

Citer

Damien Graux, Louis Jachiet, Pierre Genevès, Nabil Layaïda. SPARQLGX: Efficient Distributed Evaluation of SPARQL with Apache Spark. The 15th International Semantic Web Conference, Oct 2016, Kobe, Japan. ⟨10.1007/978-3-319-46547-0_9⟩. ⟨hal-01344915⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-RENNES1 UGA CNRS INRIA IRISA LIG INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES UR1-MATH-NUM LIG_SIDCH

600 Consultations

1455 Téléchargements

SPARQLGX: Efficient Distributed Evaluation of SPARQL with Apache Spark

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager