让世界模型学会理解3D空间中的动态轨迹,提升机器人动作精度。
4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

- 通过轨迹场对齐,让模型在时空中感知运动变化。
- 在多种基线模型上实现空间理解与执行精度显著提升。
- 适合需要精准动作规划的机器人场景应用。
基于近期世界模型进展,世界动作模型(WAMs)联合建模视频预测与动作生成。然而,它们通常在2D像素空间表示视频,与机器人执行动作的3D空间存在表征鸿沟。现有3D方法引入3D信息,但未能充分捕捉3D结构的动力学特性。本文提出4D-WAM,一种模型无关的训练策略,通过表征对齐将3D轨迹场中的时空知识注入WAMs。为此,引入两个互补目标:1)运动对齐,对齐相邻帧间的时序特征变化,鼓励模型在训练中建立局部4D感知;2)终点对齐,通过最小化源帧与目标帧间注意力相似度分布的差距,引导模型从起点推断终点。两者共同提供局部运动监督与长程目标引导,使WAMs学习到轨迹级时空表征。跨不同基线模型的分布内与分布外实验表明,该方法显著提升了空间理解、执行精度、鲁棒性、泛化能力与通用性。
原文摘要 · Abstract (English)
Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel space, creating a representation gap with 3D space in which robotic actions are executed. Recent 3D approaches introduce 3D information, but fail to fully exploit the dynamics of 3D structures. In this work, we propose 4D-WAM, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment. To this end, we introduce two complementary objectives: 1) motion alignment, which aligns temporal feature variations across adjacent frames and encourages the model to build local 4D awareness during training, and 2) destination alignment, which guides the model to infer the final destination from the source frame by minimizing the gap between their attention-like similarity distributions. Together, these objectives provide both local motion supervision and long-horizon goal guidance, enabling WAMs to learn trajectory-level spatiotemporal representations. Extensive in-distribution and out-of-distribution experiments across different base models demonstrate the model's improvements in spatial understanding, execution precision, robustness, generalization, and versatility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。