arXiv:2608.22419cs.ROcs.CV2026-08

通过简单遮蔽模态通道,提升双臂机器人视觉语言动作模型的鲁棒性。

Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking

论文配图:Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking
图 1 · 摘自论文原文
  • 训练时随机遮蔽部分模态通道,让策略学会依赖可靠信息。
  • 在10个双臂任务中成功率达21.7%,真实场景下提升超30%。
  • 无需修改架构或大量预训练,适合实际部署的机器人系统。

基于查询的视觉-语言-动作(VLA)模型具备低延迟推理优势,适用于双臂机器人操作,但复杂双臂任务中仍存在动作不连续和执行失败问题。我们观察到,多视角与语言融合不稳定是原因之一,常伴随注意力扩散至干扰区域。为此,提出仅需训练阶段的模态遮蔽机制(M3),无需架构改动或大规模机器人预训练。M3在训练中随机遮蔽部分模态通道,使策略暴露于可控的不完整观测,从而减少对干扰线索的依赖,增强对可靠信息的利用。在RoboTwin 2.0的10个双臂任务及3个长时序真实世界任务上评估,相比基线适配器(Adapter),M3在纯净设置下平均成功率提升21.7%,在Clean2Rand设置(清洁演示训练,随机场景测试)下提升11.4%,真实世界全任务成功率平均提升超过30%。结果表明,结构化训练遮蔽是提升查询式VLA策略鲁棒性的有效方法。

原文摘要 · Abstract (English)

Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that they can still exhibit discontinuous actions and execution failures in complex dual-arm tasks. We hypothesize that unstable multi-view and language fusion is one contributing factor in these failures, often coinciding with attention spreading to distracting regions. To improve robustness, we introduce the Modality Masking Mechanism (M3), an embarrassingly simple, training-only strategy that requires no architectural changes or large-scale robot pretraining. M3 stochastically masks subsets of modality channels during training, exposing the policy to controlled partial observations and encouraging it to rely less on distracting cues and more on evidence that remains reliable. We evaluate M3 on ten bimanual tasks from RoboTwin 2.0 and on three long-horizon real-world tasks. Compared with the Adapter baseline, M3 improves average success by 21.7% in the Clean setting and 11.4% in Clean2Rand, where policies are trained on clean demonstrations and evaluated on randomized scenes, while also improving averaged real-world full-task success by over 30%. These results suggest that structured training-time masking is a practical way to improve the robustness of query-based VLA policies for bimanual manipulation.

双臂操作视觉语言动作鲁棒性训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。