G3AN: Disentangling Appearance and Motion for Video Generation

Yaohui Wang; Piotr Bilinski; Francois F Bremond; Antitza Dantcheva

Communication Dans Un Congrès Année : 2020

G3AN: Disentangling Appearance and Motion for Video Generation

(1, 2) , (3) , (1, 2) , (1, 2)

1
2
3

Yaohui Wang

Fonction : Auteur
PersonId : 1058507

Spatio-Temporal Activity Recognition Systems

Université Côte d'Azur

Piotr Bilinski

Fonction : Auteur
PersonId : 1079421

University of Warsaw

Francois F Bremond

Fonction : Auteur
PersonId : 20805
IdHAL : francois-bremond
ORCID : 0000-0003-2988-2142
IdRef : 138919046

Spatio-Temporal Activity Recognition Systems

Université Côte d'Azur

Antitza Dantcheva

Fonction : Auteur
PersonId : 997689

Spatio-Temporal Activity Recognition Systems

Université Côte d'Azur

Résumé

Creating realistic human videos entails the challenge of being able to simultaneously generate both appearance, as well as motion. To tackle this challenge, we introduce G 3 AN, a novel spatio-temporal generative model, which seeks to capture the distribution of high dimensional video data and to model appearance and motion in disentangled manner. The latter is achieved by decomposing appearance and motion in a three-stream Generator, where the main stream aims to model spatio-temporal consistency, whereas the two auxiliary streams augment the main stream with multi-scale appearance and motion features, respectively. An extensive quantitative and qualitative analysis shows that our model systematically and significantly out-performs state-of-the-art methods on the facial expression datasets MUG and UvA-NEMO, as well as the Weizmann and UCF101 datasets on human action. Additional analysis on the learned latent representations confirms the successful decomposition of appearance and motion. Source code and pre-trained models are publicly available (https://wyhsirius.github.io/G3AN/)

Domaines

Vision par ordinateur et reconnaissance de formes [cs.CV]

Fichier principal

yaohuiCVPR2020.pdf (3.27 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Antitza Dantcheva : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-02969849

Soumis le : vendredi 16 octobre 2020-22:44:56

Dernière modification le : lundi 26 février 2024-11:22:14

Archivage à long terme le : dimanche 17 janvier 2021-23:38:39

Dates et versions

hal-02969849 , version 1 (16-10-2020)

Identifiants

HAL Id : hal-02969849 , version 1

Citer

Yaohui Wang, Piotr Bilinski, Francois F Bremond, Antitza Dantcheva. G3AN: Disentangling Appearance and Motion for Video Generation. CVPR 2020 - IEEE Conference on Computer Vision and Pattern Recognition, Jun 2020, Seattle / Virtual, United States. ⟨hal-02969849⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INRIA INRIA2 UNIV-COTEDAZUR OPAL 3IA-COTEDAZUR ANR

72 Consultations

131 Téléchargements

G3AN: Disentangling Appearance and Motion for Video Generation

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager