arXiv:2510.01642cs.RO2025-10被引 29

让机器人学会识别并自动恢复执行失败,提升视觉语言动作模型的鲁棒性。

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

  • 自动生成多样失败场景与可执行恢复动作,构建失败-应对数据对。
  • 在ManiSkill上使3个主流模型平均性能提升22.6%。
  • 支持不同物体、相机视角和机械臂形态的泛化,适配多种任务。

近期机器人操作研究将低层控制整合进视觉语言模型(VLM),发展为视觉语言动作(VLA)模型。尽管先进VLA在下游任务中表现优异,依赖大规模众包机器人训练数据,但执行过程中仍不可避免发生故障。如何让机器人自主推理并从突发故障中恢复仍是关键挑战。现有数据集多仅提供真实轨迹,缺乏恢复能力;少数含故障检测的数据也仅提供文本解释,难以直接用于VLA模型。为此,我们提出FailSafe,一种自动生成失败案例并配以可执行恢复动作的新系统。该系统可适配支持运动规划的模拟器,实现失败-动作数据的规模化生成。为验证效果,我们微调LLaVA-OneVision-7B构建FailSafe-VLM。实验表明,FailSafe-VLM能有效帮助机械臂检测并恢复潜在故障,在ManiSkill多个任务中使三款主流VLA模型(Pi-0-FAST、OpenVLA、OpenVLA-OFT)平均性能提升22.6%。此外,该模型具备跨空间配置、相机视角、物体与机械臂形态的泛化能力。

原文摘要 · Abstract (English)

Recent advances in robotic manipulation have integrated low-level robotic control into Vision-Language Models (VLMs), extending them into Vision-Language-Action (VLA) models. Although state-of-the-art VLAs achieve strong performance in downstream robotic applications, supported by large-scale crowd-sourced robot training data, they still inevitably encounter failures during execution. Enabling robots to reason and recover from unpredictable and abrupt failures remains a critical challenge. Existing robotic manipulation datasets, collected in either simulation or the real world, primarily provide only ground-truth trajectories, leaving robots unable to recover once failures occur. Moreover, the few datasets that address failure detection typically offer only textual explanations, which are difficult to utilize directly in VLA models. To address this gap, we introduce FailSafe, a novel failure generation and recovery system that automatically produces diverse failure cases paired with executable recovery actions. FailSafe can be easily adapted to a wide range of manipulation tasks in simulators with motion planning support, enabling scalable creation of failure-action data. To demonstrate its effectiveness, we fine-tune LLaVA-OneVision-7B (LLaVA-OV-7B) to build FailSafe-VLM. Experimental results show that FailSafe-VLM successfully helps robotic arms detect and recover from potential failures, improving the performance of three state-of-the-art VLA models (Pi-0-FAST, OpenVLA, OpenVLA-OFT) by up to 22.6% on average across several tasks in ManiSkill. Furthermore, FailSafe-VLM could generalize across different spatial configurations, camera viewpoints, object and robotic embodiments.

机器人故障恢复视觉语言动作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。