How to measure the topological quality of protein grammars?

Witold Dyrka; François Coste; Olgierd Unold; Lukasz Culer; Agnieszka Kaczmarek

Communication Dans Un Congrès Année : 2016

How to measure the topological quality of protein grammars?

(1) , (2) , (1) , (1) , (1)

1
2

Witold Dyrka

Fonction : Auteur
PersonId : 968113

Wroclaw University of Science and Technology

François Coste

Fonction : Auteur
PersonId : 9592
IdHAL : francois-coste
ORCID : 0000-0001-9134-6557
IdRef : 133160203

Dynamics, Logics and Inference for biological Systems and Sequences

Olgierd Unold

Fonction : Auteur

Wroclaw University of Science and Technology

Lukasz Culer

Fonction : Auteur

Wroclaw University of Science and Technology

Agnieszka Kaczmarek

Fonction : Auteur

Wroclaw University of Science and Technology

Résumé

Motivation. Context-free (CF) and context-sensitive (CS) formal grammars are often regarded as more appropriate to model proteins than regular level models such as finite state automata and Hidden Markov Models (HMM). In theory, the claim is well founded in the fact that many biologically relevant interactions between residues of protein sequences have a character of nested or crossed dependencies. In practice, there is hardly any evidence that grammars of higher expressiveness have an edge over old good HMMs in typical applications including recognition and classification of protein sequences. This is in contrast to RNA modeling, where CFG power some of the most successful tools. There have been proposed several explanations of this phenomenon. On the biology side, one difficulty is that interactions in proteins are often less specific and more " collective " in comparison to RNA. On the modeling side, a difficulty is the larger alphabet which combined with high complexity of CF and CS grammars imposes considerable trade-offs consisting on information reduction or learning sub-optimal solutions. Indeed, some studies hinted that CF level of expressiveness brought an added value in protein modeling when CF and regular grammars where implemented in the same framework (Dyrka, 2007; Dyrka et al., 2013). However, there have been no systematic study of explanatory power provided by various grammatical models. The first step to this goal is define objective criteria of such evaluation. Intuitively, a decent explanatory grammar should generate topology, or the parse tree, consistent with topology of the protein, or its secondary and/or tertiary structure. In this piece of research we build on this intuition and propose a set of measures to compare topology of the parse tree of a grammar with topology of the protein structure.

Mots clés

context-free grammar topological quality silhouette value protein grammar contact map

Domaines

Bio-Informatique, Biologie Systémique [q-bio.QM]

Fichier principal

icgi16_abs_rev.pdf (158.77 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

François Coste : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-01406331

Soumis le : jeudi 1 décembre 2016-09:42:34

Dernière modification le : vendredi 24 mars 2023-14:53:03

Archivage à long terme le : lundi 20 mars 2017-20:21:05

Dates et versions

hal-01406331 , version 1 (01-12-2016)

Identifiants

HAL Id : hal-01406331 , version 1
ARXIV : 1611.10078

Citer

Witold Dyrka, François Coste, Olgierd Unold, Lukasz Culer, Agnieszka Kaczmarek. How to measure the topological quality of protein grammars?. ICGI 2016 - 13th International Conference on Grammatical Inference, Oct 2016, Delft, Netherlands. ⟨hal-01406331⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSTITUT-TELECOM UNIV-RENNES1 CNRS INRIA INSA-RENNES IRISA IRISA-D7 INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES UR1-MATH-NUM

174 Consultations

42 Téléchargements

How to measure the topological quality of protein grammars?

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager