用4D高斯点云建模动态物体,提升视频预测效率与精度
4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting

- 基于4D高斯溅射分离建模动态物体与静态背景
- 在KITTI-MOT上实现短时预测与历史重建,效果优于2D模型
- 适合需要高效时空建模的自动驾驶与视频生成任务
现有世界动作模型(WAM)多基于2D视觉数据,虽能生成高质量图像,但缺乏对个体物体的显式空间结构,且重复处理冗余背景。尽管点云可表示3D场景,但跨视角对齐与累积困难。本文提出4DGS-WAM,利用显式的4D高斯溅射(4DGS)表示,分别建模动态物体与静态背景。动态物体通过策略模型预测未来动作,世界模型预测其高斯点的变换;静态背景无需重生成,因其大部分内容已在过往帧中观测到。该对象中心的世界动作模型将2D观测提升至持久的4D表示,使过去观测的静态内容可在未来预测中复用,从而聚焦于动态物体演化。在KITTI-MOT数据集上的实验评估了短时预测与历史重建性能。
原文摘要 · Abstract (English)
Current world action models (WAMs) typically operate on 2D visual data. These models can achieve exceptional visual quality, but they lack explicit spatial structure for individual objects and repeatedly process redundant background content. Although point clouds can represent the world in 3D space, they can be difficult to align and accumulate across viewpoints. In this paper, we leverage an explicit 4D Gaussian Splatting (4DGS) representation that separately models dynamic objects and the static background of a scene. For dynamic objects, we use a policy model to predict future actor actions and a world model to predict transformations of their observed Gaussian splats. The static background need not be regenerated for future states, as much of it has already been observed in past frames. This forms an object-centric world action model, which we name 4DGS-WAM. It lifts 2D observations into a persistent 4D representation so that previously observed static content can be reused during future prediction. Future-state extrapolation can then focus on modeling the evolution of dynamic objects. Experiments on KITTI-MOT evaluate short-horizon prediction and past reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。