系统梳理生成式AI三大模型的演进与应用,揭示技术突破与潜在风险。
Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- 构建GAN、VAE、扩散模型的分类体系,整合主流变体与融合方法
- 总结生成质量、多样性与可控性提升的关键技术创新
- 聚焦合成内容滥用风险,提出未来研究方向
近年来,基于深度学习的生成模型,尤其是生成对抗网络(GANs)、变分自编码器(VAEs)和扩散模型(DMs),在图像与视频合成等多领域生成高质量、多样化内容方面发挥了重要作用,推动了广泛应用并引发公众强烈兴趣。随着技术快速演进,研究文献激增、应用范围扩展以及未解决的技术挑战使得保持前沿追踪日益困难。为此,本文提出一个全面的分类体系,系统组织相关文献,为理解GAN、VAE与DM的发展脉络提供统一框架,涵盖其众多变体与组合策略。我们重点阐述提升生成内容质量、多样性和可控性的关键创新,反映生成式人工智能的拓展潜力。除技术进展外,还探讨日益凸显的伦理问题,包括滥用风险与合成媒体的社会影响。最后,指出持续存在的挑战,并提出未来研究方向,为该快速发展的领域提供结构化且前瞻性的视角。
原文摘要 · Abstract (English)
In recent years, deep learning based generative models, particularly Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Diffusion Models (DMs), have been instrumental in in generating diverse, high-quality content across various domains, such as image and video synthesis. This capability has led to widespread adoption of these models and has captured strong public interest. As they continue to advance at a rapid pace, the growing volume of research, expanding application areas, and unresolved technical challenges make it increasingly difficult to stay current. To address this need, this survey introduces a comprehensive taxonomy that organizes the literature and provides a cohesive framework for understanding the development of GANs, VAEs, and DMs, including their many variants and combined approaches. We highlight key innovations that have improved the quality, diversity, and controllability of generated outputs, reflecting the expanding potential of generative artificial intelligence. In addition to summarizing technical progress, we examine rising ethical concerns, including the risks of misuse and the broader societal impact of synthetic media. Finally, we outline persistent challenges and propose future research directions, offering a structured and forward looking perspective for researchers in this fast evolving field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。