首页 / 资料库 / 文献详情

VQ-VDM: Video Diffusion Models with 3D VQGAN

Ryota Kaji‪Keiji Yanai‬

2023Computer Science被引 2开放获取

出版方页面 →

摘要

In recent years, deep generative models have achieved impressive performance such as realizing image generation that is indistinguishable from real images. Particularly, Latent Diffusion Models, one of the image generation models, have had a significant impact on society. Therefore, video generation is attracting attention as the next modality. However, video generation is more challenging than image generation due to the consideration of temporal consistency and the increase in computational complexity, since a video is a sequence of multiple frames. In this study, we propose a video generation model based on diffusion models employing 3D VQGAN, which is called VQ-VDM. The proposed model is about nine times faster than the Video Diffusion Models which directly generate videos, since our model generates a latent representation which is decoded into a video by a VQGAN decoder. Moreover, our model can generate higher quality video than prior video generation methods exclude state-of-the-art method.

引用本文(GB/T 7714)

Ryota Kaji, ‪Keiji Yanai‬. VQ-VDM: Video Diffusion Models with 3D VQGAN[J]. 未知来源, 2023.

引文网络

参考文献与被引分析加载中…

DOI:https://doi.org/10.1145/3595916.3626363

本站仅收录题录与摘要供学习参考,全文版权归属出版方;如有侵权请联系我们删除。