arXiv:2609.05178cs.RO2026-09

新基准测试让机器人学会失败后自我修复,突破纯成功率的虚假繁荣。

LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models

论文配图:LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models
图 1 · 摘自论文原文
  • 基于LIBERO构建1000+真实失败场景,分四层恢复能力评测
  • 首次系统评估机器人在失败后的空间理解与环境适应能力
  • 适合关注机器人鲁棒性与真实场景部署的研究者

视觉-语言-动作(VLA)或世界动作(WAM)模型在机器人操作中表现优异,现有最佳方法在LIBERO基准上成功率接近100%,看似已具备实际部署能力。然而,理想条件下高成功率可能掩盖真实世界的脆弱性:现有基准仅评估从预设初始状态的任务完成情况,而现实交互中必然出现抓取失败、碰撞和物体误移等故障。机器人不仅需成功执行任务,还需识别并恢复失败以继续作业。当前这一能力尚未被充分衡量,暴露了基准性能与真实可靠性之间的关键差距。为此,我们提出LIBERO-Recover基准,一个大规模的机器人操作失败恢复评测体系。基于LIBERO,我们收集了先进具身模型的真实执行失败数据,构建了超过1000个跨四个恢复层级的场景:(1)动作重试,(2)动作调整,(3)物体状态恢复,(4)环境恢复。评估四大核心能力:空间理解、物体结构推理、交互理解与拓扑推理。作为首个大规模具身失败恢复基准,LIBERO-Recover将评估重点从“能否成功”转向“能否恢复”,推动更鲁棒、更具泛化性的具身智能体发展。项目主页:https://liulin815.github.io/LIBERO-Recovery/

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100\% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing benchmarks can be misleading: success under ideal conditions does not imply real world robustness. Existing benchmarks primarily evaluate task completion from predefined initial states, while real world interactions inevitably involve failures such as failed grasps, collisions, and unintended object movements. A robot must therefore not only execute tasks successfully, but also recognize and recover from failures to continue the task. Yet this capability remains largely unmeasured, revealing a critical gap between benchmark performance and real world reliability. To address this gap, we introduce LIBERO-Recover Benchmark, a large scale benchmark for failure recovery in robotic manipulation. Built upon LIBERO, we collect real execution failures from SOTA embodied models and construct 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery. We evaluate four core capabilities: spatial understanding, object structure reasoning, interaction understanding, and topological reasoning. As the first large-scale benchmark for embodied failure recovery, LIBERO-Recover shifts evaluation from \emph{Can the robot succeed?''} to \emph{Can the robot recover after failure?''}, promoting robust and generalizable embodied agents. The project will be avaible in \textcolor{blue}{https://liulin815.github.io/LIBERO-Recovery/}.

机器人失败恢复具身智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。