用定制架构与损失函数,让语义图生成连贯视频更真实。
SVS-GAN: Leveraging GANs for Semantic Video Synthesis
- 三重金字塔生成器+SPADE块,精准控制语义到图像的转换。
- 引入OASIS损失,通过分割网络提升时序一致性与视觉质量。
- 在Cityscapes和KITTI-360上超越现有最佳模型,适合视频生成研究者。
近年来,生成对抗网络(GAN)和扩散模型在语义图像合成(SIS)领域受到广泛关注。该领域已出现针对此任务设计的专用损失函数,区别于通用的图像到图像(I2I)翻译方法。本文首次正式提出语义视频合成(SVS)——从语义图生成时间连贯、逼真的图像序列。尽管已有部分工作探索该方向,但多数依赖通用视频到视频翻译的损失函数,或需额外数据以保证时序一致性。为此,本文提出SVS-GAN框架,专为SVS设计,包含定制架构与损失函数。其采用三重金字塔生成器,结合SPADE模块;图像判别器则基于U-Net结构,用于执行语义分割以支持OASIS损失。通过架构与目标函数的协同设计,该框架有效弥合了SIS与SVS之间的差距,在Cityscapes和KITTI-360数据集上优于当前最优模型。
原文摘要 · Abstract (English)
In recent years, there has been a growing interest in Semantic Image Synthesis (SIS) through the use of Generative Adversarial Networks (GANs) and diffusion models. This field has seen innovations such as the implementation of specialized loss functions tailored for this task, diverging from the more general approaches in Image-to-Image (I2I) translation. While the concept of Semantic Video Synthesis (SVS)$\unicode{x2013}$the generation of temporally coherent, realistic sequences of images from semantic maps$\unicode{x2013}$is newly formalized in this paper, some existing methods have already explored aspects of this field. Most of these approaches rely on generic loss functions designed for video-to-video translation or require additional data to achieve temporal coherence. In this paper, we introduce the SVS-GAN, a framework specifically designed for SVS, featuring a custom architecture and loss functions. Our approach includes a triple-pyramid generator that utilizes SPADE blocks. Additionally, we employ a U-Net-based network for the image discriminator, which performs semantic segmentation for the OASIS loss. Through this combination of tailored architecture and objective engineering, our framework aims to bridge the existing gap between SIS and SVS, outperforming current state-of-the-art models on datasets like Cityscapes and KITTI-360.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。