arXiv:2603.24584cs.CVcs.RO2026-03被引 2

让机器人在杂乱环境中更准抓目标,避免误抓错物。

TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models

  • 推理时通过对比有无物体的视觉输入,生成纠正信号
  • 在多个基准上减少误抓和近似错误,提升鲁棒性
  • 无需修改模型,适配现有视觉语言动作模型

视觉-语言-动作(VLA)策略在将语言指令与视觉观测映射为机器人动作方面取得了显著进展,但在存在干扰物的杂乱场景中其可靠性下降。分析失败案例发现,许多错误并非源于不可行的动作,而是实例级定位失败:策略常生成看似合理的抓取轨迹,但轻微偏离目标或抓到错误的物体实例。为此,我们提出一种名为TAG(目标无关引导)的简单推理时引导机制,显式降低VLA策略中干扰物和外观引起的偏差。受无分类器引导(CFG)启发,TAG对比原始观测与物体擦除后的观测下策略预测结果,利用二者差异作为残差引导信号,增强决策过程中物体证据的影响。TAG无需修改策略结构,可几乎零成本集成至现有VLA策略。我们在LIBERO、LIBERO-Plus和VLABench等标准操作基准上评估,结果表明其在杂乱环境下持续提升鲁棒性,显著减少近似失误与错误物体执行。

原文摘要 · Abstract (English)

Vision--Language--Action (VLA) policies have shown strong progress in mapping language instructions and visual observations to robotic actions, yet their reliability degrades in cluttered scenes with distractors. By analyzing failure cases, we find that many errors do not arise from infeasible motions, but from instance-level grounding failures: the policy often produces a plausible grasp trajectory that lands slightly off-target or even on the wrong object instance. To address this issue, we propose TAG (Target-Agnostic Guidance), a simple inference-time guidance mechanism that explicitly reduces distractor- and appearance-induced bias in VLA policies. Inspired by classifier-free guidance (CFG), TAG contrasts policy predictions under the original observation and an object-erased observation, and uses their difference as a residual steering signal that strengthens the influence of object evidence in the decision process. TAG does not require modifying the policy architecture and can be integrated with existing VLA policies with minimal training and inference changes. We evaluate TAG on standard manipulation benchmarks, including LIBERO, LIBERO-Plus, and VLABench, where it consistently improves robustness under clutter and reduces near-miss and wrong-object executions.

机器人控制视觉导航多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。