系统梳理3D/4D世界建模的最新进展与标准框架
3D and 4D World Modeling: A Survey
- 首次构建3D/4D世界建模的完整分类体系
- 涵盖视频、体素、激光雷达三类主流方法
- 提供数据集与评估指标的统一参考
世界建模已成为人工智能研究的核心,使智能体能够理解、表示并预测其所处的动态环境。尽管以往工作多聚焦于2D图像与视频的生成方法,却忽视了日益增长的基于原生3D与4D表示(如RGB-D图像、体素网格、激光雷达点云)的大规模场景建模研究。同时,
原文摘要 · Abstract (English)
World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphasizes generative methods for 2D image and video data, they overlook the rapidly growing body of work that leverages native 3D and 4D representations such as RGB-D imagery, occupancy grids, and LiDAR point clouds for large-scale scene modeling. At the same time, the absence of a standardized definition and taxonomy for "world models" has led to fragmented and sometimes inconsistent claims in the literature. This survey addresses these gaps by presenting the first comprehensive review explicitly dedicated to 3D and 4D world modeling and generation. We establish precise definitions, introduce a structured taxonomy spanning video-based (VideoGen), occupancy-based (OccGen), and LiDAR-based (LiDARGen) approaches, and systematically summarize datasets and evaluation metrics tailored to 3D/4D settings. We further discuss practical applications, identify open challenges, and highlight promising research directions, aiming to provide a coherent and foundational reference for advancing the field. A systematic summary of existing literature is available at https://github.com/worldbench/awesome-3d-4d-world-models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。