arXiv:2606.21088cs.RO2026-06被引 1

通过跨模态对齐提升机器人在复杂环境中的操作泛化能力。

MV-WAM: Manifold-Aware World Action Model with Value Augmentation

论文配图:MV-WAM: Manifold-Aware World Action Model with Value Augmentation
图 1 · 摘自论文原文
  • 设计跨模态因果掩码,将动作与视频预测、价值函数关联。
  • 在模拟环境中实现55.7%成功率,比最强基线高29.3%。
  • 支持真实双臂机器人执行多类任务,成功率达77.5%。

在具身机器人领域,实现在多样环境下的鲁棒且泛化性强的操作仍是一大挑战。近期的世界动作模型虽在域内表现优异,但其性能在分布外场景中并未同比提升。我们归因于视觉与动作模态之间的结构不匹配,其内在异质流形导致联合优化在分布偏移下过度损害动作鲁棒性。为此,我们提出MV-WAM,一种端到端框架,联合建模视觉预测、动作生成与价值估计,利用视频先验在训练与推理阶段增强动作泛化。核心在于跨模态因果掩码,层级地将动作锚定于预测视频帧,并将价值函数标记同时嵌入双模态。为缩小泛化差距,MV-WAM采用流形感知优化策略,显式考虑模态间结构异质性。最后,引入进展-价值调节机制,估计任务完成度并检测预测帧与生成动作间的错位,使策略能自主识别执行偏差并通过价值引导回滚恢复。在RoboTwin仿真环境中,MV-WAM在无随机动作监督的随机场景下达到55.7%的平均成功率,超越最强基线29.3%;在双臂机器人上对四类不同难度的真实任务实现77.5%的平均成功率。结果表明,流形感知的跨模态对齐对鲁棒策略泛化至关重要,为可部署的机器人操作提供新路径。

原文摘要 · Abstract (English)

Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domain performance, yet their gains do not extend proportionally to out-of-distribution scenarios. We attribute this to a structural mismatch between visual and action modalities, whose intrinsically heterogeneous manifolds cause joint optimization to disproportionately degrade action robustness under distribution shift. To address this, we propose MV-WAM, a novel end-to-end framework that jointly models visual prediction, action generation, and value estimation designed to effectively leverage video priors during both training and inference for enhanced action generalization. Key to this unification is a cross-modality causal mask that hierarchically grounds actions in predicted video frames and value function tokens in both modalities. To further narrow the generalization gap, MV-WAM adopts a manifold-aware optimization scheme that explicitly accounts for the structural heterogeneity across modalities. Finally, MV-WAM introduces a progress-value regulation mechanism that estimates task completion and detects misalignment between predicted frames and generated actions, enabling the policy to autonomously identify execution deviations and recover through value-guided rollback. On the RoboTwin simulation, MV-WAM achieves a 55.7% mean success rate on random scenarios without any randomized action supervision, outperforming the strongest baseline by 29.3%. MV-WAM achieves a 77.5% mean success rate across four real-world tasks of varying difficulty on a dual-arm robot. Our results demonstrate that manifold-aware cross-modal alignment is essential for robust policy generalization, offering a path toward deployable robotic manipulation.

机器人操作跨模态对齐泛化能力强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。