arXiv:2604.11135cs.ROcs.LG2026-04被引 10

通过空间价值图显式建模操作意图,提升机器人控制的泛化能力。

AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

论文配图:AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps
图 1 · 摘自论文原文
  • 用空间价值图作为视觉与动作间的桥梁,显式表达交互位置和任务意图。
  • 在RoboTwin 2.0上达到94.0%成功率,长时序和接触敏感任务提升显著。
  • 适合需要强泛化能力的机器人控制研究,尤其关注视觉-动作对齐问题。

预训练视频生成模型为机器人控制提供了强大先验,但现有统一世界动作模型仍需大量机器人特定训练才能生成可靠动作。我们归因于结构不匹配:视频模型捕捉场景演化,而动作生成需明确推理交互位置与底层操作意图。为此,提出AIM——一种意图感知的统一世界动作模型,通过显式空间接口弥补该鸿沟。不同于直接从未来视觉表征解码动作,AIM预测对齐的空间价值图,编码任务相关的交互结构,实现面向控制的未来动态抽象。基于预训练视频生成模型,AIM在共享混合注意力架构中联合建模未来观测与价值图。采用意图因果注意力,仅通过价值表示将未来信息传递至动作分支。进一步设计自蒸馏强化学习阶段,冻结视频与价值分支,仅优化动作头,使用投影价值图响应生成密集奖励,并结合稀疏任务信号。为支持训练与评估,构建包含3万条操作轨迹的仿真数据集,含多视角同步观测、动作与价值图标注。在RoboTwin 2.0基准测试中,AIM取得94.0%平均成功率,显著优于先前统一世界动作基线。尤其在长时序与接触敏感任务中优势更明显,验证了显式空间-意图建模在连接视觉世界建模与机器人控制中的有效性。

原文摘要 · Abstract (English)

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a structural mismatch: while video models capture how scenes evolve, action generation requires explicit reasoning about where to interact and the underlying manipulation intent. We introduce AIM, an intent-aware unified world action model that bridges this gap via an explicit spatial interface. Instead of decoding actions directly from future visual representations, AIM predicts an aligned spatial value map that encodes task-relevant interaction structure, enabling a control-oriented abstraction of future dynamics. Built on a pretrained video generation model, AIM jointly models future observations and value maps within a shared mixture-of-transformers architecture. It employs intent-causal attention to route future information to the action branch exclusively through the value representation. We further propose a self-distillation reinforcement learning stage that freezes the video and value branches and optimizes only the action head using dense rewards derived from projected value-map responses together with sparse task-level signals. To support training and evaluation, we construct a simulation dataset of 30K manipulation trajectories with synchronized multi-view observations, actions, and value-map annotations. Experiments on RoboTwin 2.0 benchmark show that AIM achieves a 94.0% average success rate, significantly outperforming prior unified world action baselines. Notably, the improvement is more pronounced in long-horizon and contact-sensitive manipulation tasks, demonstrating the effectiveness of explicit spatial-intent modeling as a bridge between visual world modeling and robot control.

机器人控制空间建模意图感知视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。