Simka: fast kmer-based method for estimating the similarity between numerous metagenomic datasets

Comparative metagenomics aims to provide high-level information based on DNA material sequenced from different environments. The purpose is mainly to estimate proximity between two or more environmental sites at the genomic level. One way to estimate similarity is to count the number of similar DNA fragments. From a computational point of view, the problem is thus to calculate the intersections between datasets of reads. Resorting to traditional methods such as all-versus-all sequence alignment is not possible on current metagenomic projects. For instance, the Tara Oceans project involves hundreds of datasets of more than 100M reads each. Maillet et al. defined the following heuristic in their method called Commet[1]. Two reads are considered similar if they share t non-overlapping kmers (words of length k). This method is currently the fastest but still does not scale on Tara Oceans samples. To tackle this issue, we introduce a new similarity function between two datasets, called Simka, based on their amount of shared kmers. To scale on large metagenomic projects, we use a new technique which is able to count the kmers of N datasets simultaneously. This method also offers new possibilities such as filtering low frequency kmers which potentially contain sequencing errors. Simka was tested and compared to Commet on 21 Tara Oceans samples. This shows that our kmer-based similarity function is very close to the read-based one of Commet. Regarding sample proximity, both methods identify the same clusters of datasets. Commet required a few weeks to compute all the intersections whereas Simka took only 4 hours.

Mots clés

comparative metagenomic numerous datasets proximity similarity intersection kmer Tara Oceans ecology

Domaines

Bio-informatique [q-bio.QM]

Fichier principal

poster5.pdf (10.48 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Gaëtan BENOIT : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-01180603

Soumis le : lundi 27 juillet 2015-15:32:57

Dernière modification le : vendredi 24 mars 2023-14:53:00

Archivage à long terme le : mercredi 28 octobre 2015-11:02:15

Dates et versions

hal-01180603 , version 1 (27-07-2015)

Identifiants

HAL Id : hal-01180603 , version 1
DOI : 10.1093/Bioinforma-cs/btu406

Citer

Gaëtan Benoit, Pierre Peterlongo, Dominique Lavenier, Claire Lemaitre. Simka: fast kmer-based method for estimating the similarity between numerous metagenomic datasets. JOBIM 2015, Jul 2015, Clermont-Ferrand, France. , ⟨10.1093/Bioinforma-cs/btu406⟩. ⟨hal-01180603⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSTITUT-TELECOM UNIV-RENNES1 CNRS INRIA INSA-RENNES IRISA CENTRALESUPELEC IRISA-D7 INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES UR1-MATH-NUM

1162 Consultations

373 Téléchargements