梳理十年来图像生成模型的技术演进与关键突破。
Image Generation Models: A Technical History
- 按技术路线系统解析VAE、GAN、扩散模型等核心架构
- 归纳各类模型的训练机制与常见缺陷,涵盖视频生成新进展
- 聚焦生成模型的伦理风险与鲁棒性部署方案
过去十年间,图像生成技术迅猛发展,但相关文献在不同模型与应用领域间显得零散。本文旨在全面综述突破性图像生成模型,包括变分自编码器(VAEs)、生成对抗网络(GANs)、归一化流、自回归与基于Transformer的生成器,以及基于扩散的方法。我们对每类模型进行详尽的技术剖析,涵盖其核心目标、架构组件及算法训练步骤。针对每种模型类型,分析优化策略、常见失败模式与局限性。此外,还回顾了视频生成领域的最新进展,阐述从静态图像到高质量视频实现的关键研究工作。最后,探讨生成模型稳健性与负责任部署的重要性,涵盖深度伪造风险、检测技术、生成伪影及水印等议题。
原文摘要 · Abstract (English)
Image generation has advanced rapidly over the past decade, yet the literature seems fragmented across different models and application domains. This paper aims to offer a comprehensive survey of breakthrough image generation models, including variational autoencoders (VAEs), generative adversarial networks (GANs), normalizing flows, autoregressive and transformer-based generators, and diffusion-based methods. We provide a detailed technical walkthrough of each model type, including their underlying objectives, architectural building blocks, and algorithmic training steps. For each model type, we present the optimization techniques as well as common failure modes and limitations. We also go over recent developments in video generation and present the research works that made it possible to go from still frames to high quality videos. Lastly, we cover the growing importance of robustness and responsible deployment of these models, including deepfake risks, detection, artifacts, and watermarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。