arXiv:2602.17659cs.CVcs.RO2026-02被引 18

提出新方法缓解视觉模型误读指令的问题,提升机器人执行语言命令的准确性。

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

  • 设计双分支推理机制,对比有无语言条件下的动作选择。
  • 在基准测试中语言遵循准确率提升9.7%,罕见任务成功率提高3.6%。
  • 无需额外训练或修改模型,可直接集成到现有系统中使用。

视觉-语言-动作模型(VLAs)旨在将语言指令与机器人控制对齐,但在实际中常未能忠实遵循语言。当指令缺乏强场景特定监督时,VLAs会因数据集偏差产生视觉捷径,反复执行训练中常见行为,忽略语言意图。为此,我们提出首个针对VLAs的反事实基准LIBERO-CF,通过在视觉合理的LIBERO布局下赋予替代指令来评估语言遵循能力。评估显示,反事实失败在主流VLAs中普遍存在但研究不足。我们提出反事实动作引导(CAG),一种简单有效的双分支推理机制,显式规范语言条件。CAG结合标准VLA策略与语言无关的视觉-动作(VA)模块,在动作选择阶段实现反事实比较。该设计减少对视觉捷径的依赖,提升对低观测任务的鲁棒性,且无需额外示范或修改现有架构与预训练模型。大量实验表明其可即插即用,跨多种VLAs持续提升性能。例如在LIBERO-CF上,采用无训练策略时,语言遵循准确率提升9.7%(π_{0.5}),罕见任务成功率提升3.6%;若搭配VA模型,进一步提升至15.5%和8.5%。真实世界评估中,反事实失败降低9.4%,任务成功率平均提升17.2%。

原文摘要 · Abstract (English)

Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow language. When presented with instructions that lack strong scene-specific supervision, VLAs suffer from counterfactual failures: they act based on vision shortcuts induced by dataset biases, repeatedly executing well-learned behaviors and selecting objects frequently seen during training regardless of language intent. To systematically study it, we introduce LIBERO-CF, the first counterfactual benchmark for VLAs that evaluates language following capability by assigning alternative instructions under visually plausible LIBERO layouts. Our evaluation reveals that counterfactual failures are prevalent yet underexplored across state-of-the-art VLAs. We propose Counterfactual Action Guidance (CAG), a simple yet effective dual-branch inference scheme that explicitly regularizes language conditioning in VLAs. CAG combines a standard VLA policy with a language-unconditioned Vision-Action (VA) module, enabling counterfactual comparison during action selection. This design reduces reliance on visual shortcuts, improves robustness on under-observed tasks, and requires neither additional demonstrations nor modifications to existing architectures or pretrained models. Extensive experiments demonstrate its plug-and-play integration across diverse VLAs and consistent improvements. For example, on LIBERO-CF, CAG improves $π_{0.5}$ by 9.7% in language following accuracy and 3.6% in task success on under-observed tasks using a training-free strategy, with further gains of 15.5% and 8.5%, respectively, when paired with a VA model. In real-world evaluations, CAG reduces counterfactual failures of 9.4% and improves task success by 17.2% on average.

视觉语言机器人控制反事实学习动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。