arXiv:2505.05474cs.CV2025-05IJCV综述被引 43

综述3D场景生成前沿,涵盖四类方法与未来方向

3D Scene Generation: A Survey

  • 按程序生成、神经3D、图像与视频驱动四类归纳方法
  • 整合主流数据集与评估标准,对比各方法优劣
  • 适合研究生成模型、3D视觉及具身智能的学者参考

3D场景生成旨在为沉浸式媒体、机器人、自动驾驶和具身智能等应用合成空间结构合理、语义明确且逼真的环境。早期基于规则的方法虽可扩展但多样性有限。近年来,生成对抗网络(GANs)、扩散模型等深度生成模型,以及神经辐射场(NeRF)、3D高斯等3D表示技术的发展,使真实世界场景分布的学习成为可能,显著提升了生成质量、多样性和视角一致性。近期进展如扩散模型通过将生成问题重新表述为图像或视频生成任务,弥合了3D场景合成与真实感之间的鸿沟。本综述系统梳理了最新技术,将其分为四类范式:程序化生成、基于神经3D的生成、基于图像的生成和基于视频的生成。分析其技术基础、权衡关系与代表性成果,回顾常用数据集、评估协议及下游应用。最后讨论生成能力、3D表示、数据与标注、评估等方面的挑战,并展望更高保真度、物理感知与交互式生成、统一感知-生成模型等方向。本文整理了3D场景生成的最新进展,指出了生成AI、3D视觉与具身智能交叉领域的前景。为追踪最新动态,项目主页持续更新:https://github.com/hzxie/Awesome-3D-Scene-Generation。

原文摘要 · Abstract (English)

3D scene generation seeks to synthesize spatially structured, semantically meaningful, and photorealistic environments for applications such as immersive media, robotics, autonomous driving, and embodied AI. Early methods based on procedural rules offered scalability but limited diversity. Recent advances in deep generative models (e.g., GANs, diffusion models) and 3D representations (e.g., NeRF, 3D Gaussians) have enabled the learning of real-world scene distributions, improving fidelity, diversity, and view consistency. Recent advances like diffusion models bridge 3D scene synthesis and photorealism by reframing generation as image or video synthesis problems. This survey provides a systematic overview of state-of-the-art approaches, organizing them into four paradigms: procedural generation, neural 3D-based generation, image-based generation, and video-based generation. We analyze their technical foundations, trade-offs, and representative results, and review commonly used datasets, evaluation protocols, and downstream applications. We conclude by discussing key challenges in generation capacity, 3D representation, data and annotations, and evaluation, and outline promising directions including higher fidelity, physics-aware and interactive generation, and unified perception-generation models. This review organizes recent advances in 3D scene generation and highlights promising directions at the intersection of generative AI, 3D vision, and embodied intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/hzxie/Awesome-3D-Scene-Generation.

3D生成生成模型综述具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。