arXiv:2606.02307cs.RO2026-06被引 1

主动发现视觉语言动作模型的隐藏失败点,提升评估可靠性。

FATE-VLA:Failue-aware test generation for vision-language-action models

论文配图:FATE-VLA:Failue-aware test generation for vision-language-action models
图 1 · 摘自论文原文
  • 用代理模型引导探索,聚焦高风险场景区域。
  • 在4个主流模型上发现最多29.7%更多失败,如GR00T-N1.6成功率从64.4%降至34.7%。
  • 适合机器人部署前的鲁棒性测试,尤其关注罕见但关键的失败场景。

视觉-语言-动作(VLA)模型正被广泛用作通用机器人策略,但其评估仍依赖静态基准,随机采样任务场景。在高维具身空间中,失败稀疏且聚集,静态基准会低估鲁棒性风险。本文将VLA评估重构为一个主动的失败发现任务,提出一种故障感知的测试生成方法,结合多样性驱动的探索与从观测执行中学习的代理模型,引导测试向高风险且多样化的场景区域推进。在四个最先进的VLA模型上,该方法比选定基线多发现最多29.7%的失败,并揭示了更丰富的失败模式。例如,对于GR00T-N1.6,成功率从64.4%下降至34.7%。总体而言,研究呼吁改变VLA评估范式:从对固定任务集的被动测量,转向自适应、以失败为导向的测试生成,从而在部署前揭示模型弱点的结构特征。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are increasingly used as generalist robot policies, yet their evaluation still relies largely on static benchmarks that randomly sample task scenes. In high-dimensional embodied spaces, failures are sparse and clustered, so static benchmarking can underestimate robustness risks. We reframe VLA evaluation as an active failure-discovery problem and propose a failure-aware test-generation approach that combines diversity-driven exploration with surrogate models learned from observed executions. The method steers testing toward high-risk yet diverse scene regions. Across four state-of-the-art VLA models, it uncovers substantially more failures (up to +29.7 % over selected baselines) while revealing more diverse failure modes. This mean that, for instance, in the case of GR00T-N1.6, success rate dropped from 64.4% to 34.7%. More broadly, our findings call for a shift in VLA evaluation: from passive measurement on fixed task suites to adaptive, failure-seeking test generation that exposes the structure of model weaknesses before deployment.

机器人策略故障检测评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。