arXiv:2512.01946cs.ROcs.CV2025-12被引 11

用视觉语言模型生成36000条真实故障数据,提升机器人任务失败检测能力

Guardian: Detecting Robotic Planning and Execution Errors with Vision-Language Models

  • 通过扰动成功轨迹自动合成多样化故障,构建真实分布的失败数据
  • 在3个新实机基准上达到最优性能,使任务成功率显著提升
  • 适合研究机器人可靠性、故障检测与视觉语言模型应用的学者

稳健的机器人操作依赖可靠的故障检测与恢复机制。尽管近期视觉语言模型(VLMs)在故障检测中展现潜力,但其泛化能力受限于故障数据的稀缺性与覆盖范围狭窄。为此,我们提出一个自动化框架,可在仿真与真实环境中共生成多样化的机器人规划与执行故障。该方法通过扰动成功的操作轨迹,合成反映真实故障分布的失败样本,并利用VLM生成结构化的分步推理轨迹。由此构建了GuardianFail-36k,一个基于RLBench仿真器与BridgeDataV2真实机器人数据集的大规模故障推理数据集。基于此数据集,我们训练出Guardian——一个用于统一规划与执行验证的多视角推理VLM。Guardian在三个未见的真实世界基准(RoboFail、RoboVQA及新提出的UR5-Fail)上达到当前最佳表现。当与先进的基于LLM的操纵策略结合时,其在仿真与真实部署中均持续提升任务成功率。结果表明,高质量故障推理数据的规模化对提升机器人故障检测泛化能力至关重要。代码、数据与模型详见 https://www.di.ens.fr/willow/research/guardian/。

原文摘要 · Abstract (English)

Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in robot failure detection, their generalization is severely limited by the scarcity and narrow coverage of failure data. To address this bottleneck, we propose an automatic framework for generating diverse robotic planning and execution failures across both simulated and real-world environments. Our approach perturbs successful manipulation trajectories to synthesize failures that reflect realistic failure distributions, and leverages VLMs to produce structured step-by-step reasoning traces. This yields GuardianFail-36k, a large-scale failure reasoning dataset built upon the RLBench simulator and the BridgeDataV2 real-robot dataset. Using GuardianFail-36k, we train Guardian, a multi-view reasoning VLM for unified planning and execution verification. Guardian achieves state-of-the-art performance on three unseen real-world benchmarks: RoboFail, RoboVQA, and our newly introduced UR5-Fail. When integrated with a state-of-the-art LLM-based manipulation policy, it consistently boosts task success rates in both simulation and real-world deployment. These results demonstrate that scaling high-quality failure reasoning data is critical for improving generalization in robotic failure detection. Code, Data, and Models are available at https://www.di.ens.fr/willow/research/guardian/.

机器人故障检测视觉语言模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。