Seqcrawler: biological data indexing and browsing platform. - Archive ouverte HAL Access content directly
Journal Articles BMC Bioinformatics Year : 2012

Seqcrawler: biological data indexing and browsing platform.

(1) , (2) , (2)
1
2

Abstract

ABSTRACT: BACKGROUND: Seqcrawler takes its roots in software like SRS or Lucegene. It provides an indexing platform to ease the search of data and meta-data in biological banks and it can scale to face the current flow of data. While many biological bank search tools are available on the Internet, mainly provided by large organizations to search in their data, there is a lack of free and open source solution to browse one own set of data with a flexible query system and able to scale from single computer to a cloud system. A personal index platform will help labs and bioinformaticians to search in their meta-data but also to build a larger information system with custom subsets of data. RESULTS: The software is scalable from a single computer to a cloud-based infrastructure. It has been successfully tested in a private cloud with 3 index shards (piece of index) hosting ~400 millions of sequence information (whole GenBank, UniProt, PDB and others) for a total size of 600 GB in a fault tolerant architecture (high-availability). It has also been successfully integrated with software to add extra meta-data from blast results to enhance user's result analysis. CONCLUSIONS: Seqcrawler provides a complete open source search and store solution for labs or platforms needing to manage large amount of data/meta-data with a flexible and customizable web interface. All components (search engine, visualization and data storage), though independent, share a common and coherent data system that can be queried with a simple HTTP interface. The solution scales easily and can also provide a high availability infrastructure.

Dates and versions

hal-00728279 , version 1 (05-09-2012)

Identifiers

Cite

Olivier Sallou, Anthony Bretaudeau, Aurelien Roult. Seqcrawler: biological data indexing and browsing platform.. BMC Bioinformatics, 2012, 13 (1), pp.175. ⟨10.1186/1471-2105-13-175⟩. ⟨hal-00728279⟩
134 View
1 Download

Altmetric

Share

Gmail Facebook Twitter LinkedIn More