Parallel Position Weight Matrices Algorithms

Mathieu Giraud 1, 2 Jean-Stéphane Varré 1, 2
2 BONSAI - Bioinformatics and Sequence Analysis
LIFL - Laboratoire d'Informatique Fondamentale de Lille, Inria Lille - Nord Europe
Abstract : Position Weight Matrices (PWMs) are broadly used in computational biology. The basic problems, Scan and MultipleScan, aim to find all the occurrences of a given PWM or a set of PWMs in long sequences. Some other PWM tasks share a common NP-hard subproblem, ScoreDistribution. The existing algorithms rely on the enumeration on a large set of scores or words, and they are mostly not suitable for parallelization. We propose a new algorithm, BucketScoreDistribution, that is both very efficient and suitable for parallelization. We bound the error induced by this algorithm. We realized a GPU prototype for Scan, MultipleScan and BucketScoreDistribution with the CUDA libraries, and report for the different problems speedups larger than 10× on several Nvidia cards.
Mathieu Giraud, Jean-Stéphane Varré. Parallel Position Weight Matrices Algorithms. Parallel Computing, Elsevier, 2011, 37, pp.466-478. <10.1016/j.parco.2010.10.001>. <hal-00623404>



