Skip to Main content Skip to Navigation
Conference papers

Learning URI Selection Criteria to Improve the Crawling of Linked Open Data (Extended Abstract)

Hai Huang 1, 2 Fabien Gandon 1
1 WIMMICS - Web-Instrumented Man-Machine Interactions, Communities and Semantics
CRISAM - Inria Sophia Antipolis - Méditerranée , Laboratoire I3S - SPARKS - Scalable and Pervasive softwARe and Knowledge Systems
Abstract : A Linked Data crawler performs a selection to focus on collecting linked RDF (including RDFa) data on the Web. From the perspectives of throughput and coverage, given a newly discovered and targeted URI, the key issue of Linked Data crawlers is to decide whether this URI is likely to dereference into an RDF data source and therefore it is worth downloading the representation it points to. Current solutions adopt heuristic rules to filter irrelevant URIs. But when the heuristics are too restrictive this hampers the coverage of crawling. In this paper, we propose and compare approaches to learn strategies for crawling Linked Data on the Web by predicting whether a newly discovered URI will lead to an RDF data source or not. We detail the features used in predicting the relevance and the methods we evaluated including a promising adaptation of FTRL-proximal online learning algorithm. We compare several options through extensive experiments including existing crawlers as baseline methods to evaluate their efficiency.
Document type :
Conference papers
Complete list of metadata
Contributor : Hai Huang Connect in order to contact the contributor
Submitted on : Monday, December 14, 2020 - 4:03:11 PM
Last modification on : Tuesday, January 4, 2022 - 5:45:45 AM
Long-term archiving on: : Monday, March 15, 2021 - 7:49:25 PM


Files produced by the author(s)


  • HAL Id : hal-03064912, version 1


Hai Huang, Fabien Gandon. Learning URI Selection Criteria to Improve the Crawling of Linked Open Data (Extended Abstract). IJCAI 2020 - 29th International Joint Conference on Artificial Intelligence, Jan 2021, Yokohama, Japan. ⟨hal-03064912⟩



Les métriques sont temporairement indisponibles