Skip to Main content Skip to Navigation
Conference papers

G3AN: Disentangling Appearance and Motion for Video Generation

Abstract : Creating realistic human videos entails the challenge of being able to simultaneously generate both appearance, as well as motion. To tackle this challenge, we introduce G 3 AN, a novel spatio-temporal generative model, which seeks to capture the distribution of high dimensional video data and to model appearance and motion in disentangled manner. The latter is achieved by decomposing appearance and motion in a three-stream Generator, where the main stream aims to model spatio-temporal consistency, whereas the two auxiliary streams augment the main stream with multi-scale appearance and motion features, respectively. An extensive quantitative and qualitative analysis shows that our model systematically and significantly out-performs state-of-the-art methods on the facial expression datasets MUG and UvA-NEMO, as well as the Weizmann and UCF101 datasets on human action. Additional analysis on the learned latent representations confirms the successful decomposition of appearance and motion. Source code and pre-trained models are publicly available (https://wyhsirius.github.io/G3AN/)
Document type :
Conference papers
Complete list of metadatas

Cited literature [45 references]  Display  Hide  Download

https://hal.inria.fr/hal-02969849
Contributor : Antitza Dantcheva <>
Submitted on : Friday, October 16, 2020 - 10:44:56 PM
Last modification on : Tuesday, October 20, 2020 - 3:37:24 AM

File

yaohuiCVPR2020.pdf
Files produced by the author(s)

Identifiers

  • HAL Id : hal-02969849, version 1

Citation

Yaohui Wang, Piotr Bilinski, Francois Bremond, Antitza Dantcheva. G3AN: Disentangling Appearance and Motion for Video Generation. CVPR 2020 - IEEE Conference on Computer Vision and Pattern Recognition, Jun 2020, Seattle / Virtual, United States. ⟨hal-02969849⟩

Share

Metrics

Record views

15

Files downloads

55