arXiv:2606.08881cs.ROcs.AI2026-06被引 2

在低成本机器人上评估视觉语言动作模型的失败与恢复能力。

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

论文配图:Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis
图 1 · 摘自论文原文
  • 构建SO-101平台上的标准化真实世界评测基准。
  • 发现执行不稳是主要失败原因,恢复能力因模型而异。
  • 适合关注低成本机器人鲁棒性与故障应对的研究者。

视觉语言动作(VLA)模型在机器人操作中展现出强大泛化能力,但现有评估多限于仿真或昂贵机器人平台,缺乏对低成本真实机器人鲁棒性的探索。本文提出面向低成本SO-101机器人的标准化真实世界基准,包含四个典型操作任务及统一评估协议,支持在具身不确定性下的系统性对比。基于真实遥操作演示,直接在物理平台上微调并评估$π_{0.5}$、SmolVLA、Wall-X和ACT。除传统任务成功率外,基准引入结构化失败分类、语义与执行层级失败分解及恢复感知评估指标,以刻画策略鲁棒性。实验表明,更强的预训练VLA模型总体优于模仿学习基线,但性能高度依赖任务,在低成本部署条件下仍受限。执行不稳是主要失败来源,而恢复能力在不同架构间差异显著。结果强调了超越二元成功评价的重要性,并确立SO-101作为真实低成本机器人部署环境下具身AI系统评估的实用基准。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored. We present a standardized real-world benchmark for evaluating representative VLA and imitation learning policies on the low-cost SO-101 robotic platform. The benchmark comprises four representative manipulation tasks together with unified evaluation protocols, enabling systematic comparison under embodiment uncertainty. Using real-world teleoperated demonstrations, we fine-tune and evaluate $π_{0.5}$, SmolVLA, Wall-X, and ACT directly on the physical platform. Beyond conventional task success rates, the benchmark incorporates a structured failure taxonomy, semantic- and execution-level failure decomposition, and recovery-aware evaluation metrics to characterize policy robustness. Experimental results show that stronger pretrained VLA policies generally outperform the imitation learning baseline, although performance remains highly task-dependent under low-cost robotic deployment conditions. Execution instability emerges as the dominant failure source, while recovery capability varies substantially across architectures. These results highlight the importance of failure and recovery analysis beyond binary task success and establish SO-101 as a practical benchmark for evaluating embodied AI systems under realistic low-cost robotic deployment conditions.

机器人视觉语言失败分析低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。