系统梳理4D空间智能重建的五个层级,助力理解视觉场景演化。
Reconstructing 4D Spatial Intelligence: A Survey
- 按五级递进框架分类4D重建方法:从基础属性到物理规律
- 揭示各层级关键挑战,提出未来研究方向
- 适合关注三维动态场景建模与具身AI的研究者
从视觉观测中重建4D空间智能是计算机视觉的核心挑战之一,应用广泛。该领域快速发展,现有综述难以覆盖其全貌,尤其缺乏对4D场景重建层次结构的系统分析。为此,本文提出一种新视角,将现有方法归纳为五个渐进层级:(1) 低层3D属性重建(如深度、姿态、点云图);(2) 3D场景组件重建(如物体、人体、结构);(3) 4D动态场景重建;(4) 场景组件间交互建模;(5) 物理规律与约束的融入。文章在每层级讨论关键挑战,并展望迈向更丰富4D智能的方向。项目主页持续更新:https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence。
原文摘要 · Abstract (English)
Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, where the focus is often on reconstructing fundamental visual elements, to embodied AI, which emphasizes interaction modeling and physical realism. Fueled by rapid advances in 3D representations and deep learning architectures, the field has evolved quickly, outpacing the scope of previous surveys. Additionally, existing surveys rarely offer a comprehensive analysis of the hierarchical structure of 4D scene reconstruction. To address this gap, we present a new perspective that organizes existing methods into five progressive levels of 4D spatial intelligence: (1) Level 1 -- reconstruction of low-level 3D attributes (e.g., depth, pose, and point maps); (2) Level 2 -- reconstruction of 3D scene components (e.g., objects, humans, structures); (3) Level 3 -- reconstruction of 4D dynamic scenes; (4) Level 4 -- modeling of interactions among scene components; and (5) Level 5 -- incorporation of physical laws and constraints. We conclude the survey by discussing the key challenges at each level and highlighting promising directions for advancing toward even richer levels of 4D spatial intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。