arXiv:2603.06001cs.ROcs.AI2026-03被引 8

让视觉语言模型不再‘视而不见’,用不训练的方法重校注意力

Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration

  • 提出无需训练的推理时注意力重校机制,提升语言指令权重
  • 在30个LIBERO任务上,错误执行率显著下降,基线性能不受影响
  • 适合关注机器人指令可靠性、想快速提升现有模型鲁棒性的研究者

视觉-语言-动作(VLA)模型可直接根据自然语言指令执行机器人操作,被视为通用机器人策略的基础。然而,其在分布外(OOD)指令下的可靠性尚未充分探索。本文揭示一种关键失效模式:当语言指令与场景矛盾时,VLA模型仍会执行视觉上合理但逻辑错误的动作,称为“语言盲视”。为此,我们构建ICBench诊断基准,基于LIBERO数据集,在保持视觉环境不变的前提下,注入可控的语义冲突指令。对Pi0、Pi0.5和OpenVLA OFT三类代表性VLA架构的评估显示,这些模型在逻辑不可能指令下仍频繁成功,表明其行动严重依赖视觉先验。为缓解此问题,我们提出指令引导注意力重校(IGAR),一种无需训练或修改架构的推理时机制,可动态平衡注意力分布以恢复语言指令的影响。实验表明,IGAR在30个LIBERO任务中显著降低错误执行率,同时保持原有性能;在真实Franka机械臂上的验证也证明其能有效阻止由不一致指令触发的误操作。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliability under Out-of-Distribution (OOD) instructions remains underexplored. In this paper, we reveal a critical failure mode in which VLA policies continue executing visually plausible actions even when the language instruction contradicts the scene. We refer to this phenomenon as linguistic blindness, where VLA policies prioritize visual priors over instruction semantics during action generation. To systematically analyze this issue, we introduce ICBench, a diagnostic benchmark constructed from the LIBERO dataset that probes language-action coupling by injecting controlled OOD instruction contradictions while keeping the visual environment unchanged. Evaluations on three representative VLA architectures, including Pi0, Pi0.5 and OpenVLA OFT, show that these models frequently succeed at tasks despite logically impossible instructions, revealing a strong visual bias in action generation. To mitigate this issue, we propose Instruction-Guided Attention Recalibration (IGAR), a train-free inference-time mechanism that rebalances attention distributions to restore the influence of language instructions. IGAR operates without retraining or architectural modification and can be directly applied to existing VLA models. Experiments across 30 LIBERO tasks demonstrate that IGAR substantially reduces erroneous execution under OOD contradictory instructions while preserving baseline task performance. We additionally validate the approach on a real Franka robotic arm, where IGAR effectively prevents manipulation triggered by inconsistent instructions.

机器人语言理解注意力机制鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。