arXiv:2606.12217cs.CVcs.AI2026-06被引 3

让世界模型的视觉预测更利于机器人抓取决策。

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

论文配图:Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
图 1 · 摘自论文原文
  • 用语义对齐约束视频扩散特征,使中间表示更适配动作控制。
  • 在真实任务中提升物体定位准确率与抗干扰能力。
  • 适合关注机器人视觉规划与动作生成的研究者。

世界动作模型(WAMs)通过视频生成模型预测未来场景演化,再生成控制动作,为机器人操作提供了新路径。然而我们发现:生成合理的视觉未来并不保证动作准确。分析显示,动作解码器未能聚焦于任务相关交互区域,仍受无关区域扰动影响。这揭示了表征错位问题——为视觉重建优化的隐藏状态并不天然适合低层动作控制。为此,本文提出AGRA(Action-Grounded Representation Alignment),通过将视频扩散模型的中间特征与基础视觉编码器的语义一致表征对齐,规整世界-动作接口。在真实世界操作任务上的实验表明,AGRA使模型表征更具动作导向性:动作解码器能精准聚焦关键区域,显著提升物体定位精度和功能理解能力,并增强对无关区域扰动的鲁棒性。结果表明,相较基线模型,AGRA在分布内性能和分布外泛化能力上均实现稳定提升。

原文摘要 · Abstract (English)

World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a phenomenon: generating plausible visual futures does not always guarantee the extraction of accurate actions. To diagnose this failure, we conduct action-head attention analysis and causal interventions. We find that the action decoder fails to focus on task-relevant interaction regions and remains sensitive to perturbations in task-irrelevant areas. This reveals a representation mismatch: hidden states optimized for visual reconstruction are not inherently organized in a form useful for low-level action control. In this paper, we propose AGRA, an Action-Grounded Representation Alignment objective that regularizes the world-action interface by aligning intermediate video diffusion features with spatially coherent semantic representations from a foundation visual encoder. We evaluate AGRA on real-world manipulation tasks. Experiments show that AGRA makes world model representations more action-grounded: by focusing the action decoder on the correct interaction regions, it improves object localization accuracy and affordance understanding, and makes the policy more robust to perturbations in task-irrelevant regions. As a result, AGRA consistently improves both in-distribution performance and out-of-distribution generalization over the baseline world action model.

机器人操作视觉预测动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。