用掩码统一输入与预测,让机器人更准地理解语言指令。
MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

- 用混合变压器统一处理掩码输入和预测,实现对象中心建模。
- 在多个任务上显著提升语言模糊场景下的成功率,最高达73.2%。
- 适合需要精准空间理解的机器人操控任务,尤其对新手物体有效。
世界动作模型(WAMs)通过视频预测为机器人控制提供了有前景的范式。然而,现有WAMs存在根本性的空间瓶颈:标准文本输入在杂乱场景中易引发指代歧义,而无结构的RGB预测缺乏语义基础,且受任务无关背景干扰。为此,我们提出MaskWAM,一种以对象为中心的世界动作模型。通过在统一的混合变压器(MoT)框架中联合使用掩码作为显式输入与预测目标,MaskWAM实现了稳健的策略泛化。该设计带来两大优势:(1) 预测未来掩码提供对象级语义监督,有效抑制视觉噪声,显著提升标准文本条件下的WAM性能;(2) 将此预测监督与首帧视觉提示(如目标对象掩码)结合,建立精确的空间锚点,大幅降低语言歧义。关键在于,由于WAM本质上是视觉驱动架构,直接掩码条件带来的引导力远强于文本,为操纵未知物体建立了精确且鲁棒的新范式。在LIBERO、RoboTwin及真实世界任务上的评估表明,MaskWAM在语言清晰和语言模糊任务中均显著优于基线方法。
原文摘要 · Abstract (English)
World Action Models (WAMs) present a promising paradigm for robotic control via video prediction. However, current WAMs suffer from fundamental spatial bottlenecks: standard text inputs introduce referential ambiguity in cluttered scenes, while unstructured RGB predictions lack semantic grounding and remain biased by task-irrelevant backgrounds. To overcome these limitations, we introduce MaskWAM, an object-centric world-action model. By jointly integrating masks as both explicit inputs and predictions via a unified Mixture of Transformers (MoT), MaskWAM unlocks robust policy generalization. This design provides two key benefits: (1) predicting future masks yields object-centric semantic supervision that suppresses visual noise, significantly enhancing even standard text-conditioned WAMs; and (2) coupling this predictive supervision with first-frame visual prompts, such as target object masks, establishes a precise spatial anchor that substantially reduces language ambiguity. Crucially, as WAMs are inherently vision-driven architectures, direct mask conditioning yields substantially stronger guidance than text alone, establishing a precise and robust paradigm for manipulating unseen objects. Evaluations on LIBERO, RoboTwin, and real-world tasks demonstrate that MaskWAM significantly outperforms baselines in both language-clear and language-ambiguous tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。