综述自动场景生成前沿技术,涵盖模型、数据集与挑战
Automatic Scene Generation: State-of-the-Art Techniques, Models, Datasets, Challenges, and Future Prospects
- 按VAE/GAN/Transformer/扩散模型四类梳理主流生成模型
- 总结常用数据集及评估指标,如FID、mAP等关键性能指标
- 适合想快速了解该领域进展的研究者与应用开发者
自动场景生成是机器人、娱乐、视觉表征、训练模拟、教育等领域的重要研究方向。本文综述了当前基于机器学习、深度学习、嵌入式系统和自然语言处理的自动场景生成技术。将模型分为变分自编码器(VAEs)、生成对抗网络(GANs)、Transformer 和扩散模型四类,深入分析各类子模型及其贡献。回顾了常用数据集,如 COCO-Stuff、Visual Genome、MS-COCO,这些数据集对模型训练与评估至关重要。探讨了图像转3D、文本转3D、UI/布局设计、图结构方法及交互式生成等方法。讨论了弗雷切特图像距离(FID)、KL散度、生成分数(IS)、交并比(IoU)和平均精度均值(mAP)等评估指标的应用。指出领域核心挑战:保持真实感、处理多物体复杂场景、确保对象间关系与空间布局的一致性。通过总结最新进展并识别改进方向,本综述为研究人员与实践者提供重要参考。
原文摘要 · Abstract (English)
Automatic scene generation is an essential area of research with applications in robotics, recreation, visual representation, training and simulation, education, and more. This survey provides a comprehensive review of the current state-of-the-arts in automatic scene generation, focusing on techniques that leverage machine learning, deep learning, embedded systems, and natural language processing (NLP). We categorize the models into four main types: Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), Transformers, and Diffusion Models. Each category is explored in detail, discussing various sub-models and their contributions to the field. We also review the most commonly used datasets, such as COCO-Stuff, Visual Genome, and MS-COCO, which are critical for training and evaluating these models. Methodologies for scene generation are examined, including image-to-3D conversion, text-to-3D generation, UI/layout design, graph-based methods, and interactive scene generation. Evaluation metrics such as Frechet Inception Distance (FID), Kullback-Leibler (KL) Divergence, Inception Score (IS), Intersection over Union (IoU), and Mean Average Precision (mAP) are discussed in the context of their use in assessing model performance. The survey identifies key challenges and limitations in the field, such as maintaining realism, handling complex scenes with multiple objects, and ensuring consistency in object relationships and spatial arrangements. By summarizing recent advances and pinpointing areas for improvement, this survey aims to provide a valuable resource for researchers and practitioners working on automatic scene generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。