发现并修复视觉-语言-动作模型在机器人操作中的局部失效区域
Uncovering and Mitigating Positional Blind Spots in Vision-Language-Action Models

- 通过网格扫描与似然比检验定位模型的局部失效区域
- 失效率最高达58%,针对性微调使整体失败率下降40%至85%
- 适合关注机器人可靠性与泛化能力的研究者
近期的视觉-语言-动作(VLA)模型在机器人操作中表现优异,通常以预设物体配置下的成功率作为评估标准,隐含假设工作空间内性能是均匀的。然而这一假设不成立:即使指令和其他场景因素保持不变,仅移动一个任务无关干扰物,也会在特定空间区域内显著提升失败概率,我们称其为位置盲区(PBS)。本文提出一种两阶段黑箱框架来发现并缓解PBS。在发现阶段,将工作区划分为网格,使用单侧对数似然比检验定位风险显著升高的PBS单元;在缓解阶段,基于在这些区域收集的示范数据,通过LoRA对策略进行微调,提升局部能力同时基本保持全局性能。我们在两个基准上评估了五种先进VLA策略,发现所有模型均普遍存在空间集中型的PBS,最高失败率达0.58。搜索策略平均F1得分为0.678,优于随机搜索和自适应采样基线0.268和0.178。基于发现区域的定向微调使总体失败率降低40.00%–85.19%。
原文摘要 · Abstract (English)
Recent Vision-Language-Action (VLA) models achieve promising performance in robotic manipulation, typically measured by success rates aggregated over predefined object configurations, an evaluation that implicitly assumes spatially uniform competence across the workspace. However, this assumption does not hold: even with the instruction and every other scene factor held fixed, merely relocating a task-irrelevant distractor can sharply raise the failure probability within localized, spatially coherent regions, which we term Positional Blind Spots (PBS). In this paper, we propose a two-stage black-box framework to uncover and mitigate PBS. During the uncovering stage, we grid the workspace and apply a one-sided log-likelihood-ratio test to localize PBS cells with significantly elevated risk. During the mitigation stage, we fine-tune the policy via LoRA on demonstrations collected from these PBS regions, improving competence there while largely preserving performance across the rest of the workspace. We evaluate our framework on five state-of-the-art VLA policies across two benchmarks, and find that PBS are pervasive and spatially concentrated in all of them, with failure rates up to 0.58. Our search strategy achieves an average F1-score of 0.678, outperforming random search and adaptive sampling baselines by 0.268 and 0.178, respectively. Guided by the discovered regions, targeted fine-tuning reduces the overall failure rate by 40.00%--85.19%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。