用3D高斯粒子图模型提升物体运动预测精度,解决细节丢失和轨迹漂移问题。
DyG$^2$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer

- 通过补全关键点空间细节并分离时间变化特征,增强动态建模能力。
- 在合成与真实数据集上实现更准确的轨迹预测,跨对象泛化能力强。
- 适合需要高精度运动推理的机器人交互、视频生成等场景。
从有限视觉观测中建模物体动态是实现具身交互场景下精确运动轨迹预测的基础问题。现有方法先将重构的粒子表示压缩为稀疏关键点,并使用局部约束交互建模其演化,导致细粒度局部信息丢失,跨时空尺度的区分性交互建模被掩盖,引发轨迹漂移和外观预测不准。为此,我们提出 DyG²T,一种通过空间补全与时间解耦关键点表示并建模粒子图上的多尺度交互来推断物体运动轨迹的动力学建模框架。空间上,通过聚合邻近原始粒子位置丰富每个关键点,恢复细粒度局部细节,同时显式编码关键点间的相对偏移以增强几何结构感知。时间上,引入时序解耦网络(TDN)识别潜在空间中的主导跨帧变化并放大帧间差异,生成具有时间判别性的表示,再经时序注意力聚合以捕捉帧级时间演化线索。为实现全面的交互建模,粒子图变换器利用全局注意力保留关键点间的区分性长程依赖,缓解局部约束建模引起的表征同质化,为精准轨迹预测提供稳健基础。在合成与真实世界数据集上的实验表明,DyG²T 实现了精确的动力学建模与推理,并展现出强大的跨对象与真实世界泛化能力。
原文摘要 · Abstract (English)
Modeling object dynamics from limited visual observations is a fundamental problem for enabling accurate motion trajectory prediction in embodied interaction scenarios. Existing dynamics modeling methods first compress reconstructed particle representations into sparse Key Points and model their evolution using locally constrained interactions, thereby discarding fine-grained local details and obscuring discriminative interaction modeling across spatial and temporal scales, leading to drifting trajectories and inaccurate appearance prediction. To tackle these issues, we propose DyG$^2$T, a dynamics modeling framework that infers object motion trajectories by spatially completing and temporally discriminating Key Point representations and modeling multi-scale interaction over particle graphs. Spatially, DyG$^2$T enriches each Key Point by aggregating neighboring raw particle positions to recover fine-grained local details, while explicitly encoding relative offsets among Key Points to enhance geometric structure perception. Temporally, we introduce a Temporal Disentangling Network (TDN) to identify dominant cross-frame variations in latent space and amplify inter-frame differences, yielding temporally discriminative representations that are subsequently aggregated via Temporal Attention to capture frame-wise temporal evolution cues. For comprehensive interaction modeling, a Particle Graph Transformer leverages global attention to preserve discriminative long-range dependencies among Key Points, mitigating representation homogenization induced by locality-constrained modeling and providing a robust basis for accurate trajectory prediction. Experiments on both synthetic and real-world datasets demonstrate that DyG$^2$T achieves accurate dynamics modeling and reasoning, and exhibits strong cross-object and real-world generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。