arXiv:2510.19255cs.CV2025-10被引 3

系统梳理4D生成与重建的核心表示方法,助你选对技术方案。

Advances in 4D Representation: Geometry, Motion, and Interaction

论文配图:Advances in 4D Representation: Geometry, Motion, and Interaction
图 1 · 摘自论文原文
  • 按几何、运动、交互三维度分类4D表示,聚焦代表性方法
  • 对比不同场景下各类表示的性能与挑战,提供选型指导
  • 涵盖NeRF、3DGS等主流模型,也关注长时序与结构化建模

我们对4D生成与重建这一快速发展的计算机图形学子领域进行了综述,其进展得益于神经场、几何与运动深度学习以及3D生成式人工智能(GenAI)的突破。不同于以往详尽罗列的工作,本综述从4D表示的独特视角出发,聚焦于建模随时间演化的3D几何形态及其运动与交互行为。我们选取代表性方法,分析其在不同计算条件、应用场景和数据需求下的优势与局限,旨在帮助读者理解如何为特定任务选择并定制合适的4D表示。内容围绕几何、运动、交互三大支柱展开,涵盖当前主流的神经辐射场(NeRFs)与3D高斯溅射(3DGS),也关注结构化模型与长程运动等较少研究的方向。我们还探讨了大语言模型(LLMs)与视频基础模型(VFMs)在4D应用中的角色与当前限制,并总结现有4D数据集现状及缺失,以推动该领域发展。

原文摘要 · Abstract (English)

We present a survey on 4D generation and reconstruction, a fast-evolving subfield of computer graphics whose developments have been propelled by recent advances in neural fields, geometric and motion deep learning, as well as 3D generative artificial intelligence (GenAI). While our survey is not the first of its kind, we build our coverage of the domain from a unique and distinctive perspective of 4D representations, to model 3D geometry evolving over time while exhibiting motion and interaction. Specifically, instead of offering an exhaustive enumeration of many works, we take a more selective approach by focusing on representative works to highlight both the desirable properties and ensuing challenges of each representation under different computation, application, and data scenarios. The main take-away message we aim to convey to the readers is on how to select and then customize the appropriate 4D representations for their tasks. Organizationally, we separate the 4D representations based on three key pillars: geometry, motion, and interaction. Our discourse will not only encompass the most popular representations of today, such as neural radiance fields (NeRFs) and 3D Gaussian Splatting (3DGS), but also bring attention to relatively under-explored representations in the 4D context, such as structured models and long-range motions. Throughout our survey, we will reprise the role of large language models (LLMs) and video foundational models (VFMs) in a variety of 4D applications, while steering our discussion towards their current limitations and how they can be addressed. We also provide a dedicated coverage on what 4D datasets are currently available, as well as what is lacking, in driving the subfield forward. Project page:https://mingrui-zhao.github.io/4DRep-GMI/

4D生成神经场3DGS表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。