arXiv:2509.04018cs.RO2025-09被引 24

让机器人提前预判并纠正操作失误,提升复杂任务的可靠性。

FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction

  • 用视觉语言查询评估动作可行性,动态生成纠错策略。
  • 在多个仿真和真实机器人上表现优于现有模型,零样本与微调均有效。
  • 适合需要高可靠性的长期自主任务,如工业自动化、家庭服务。

机器人操作是自动化的核心。传统感知-规划流水线在开放任务中因灵活性不足而受限,而单一端到端的视觉-语言-动作(VLA)架构虽有潜力,却缺乏失败预测与恢复机制。为此,我们提出FPC-VLA,一种融合VLA与监督模块的双模型框架。该监督模块通过视觉语言查询评估动作可行性,并在风险出现时生成纠正策略,训练无需人工标注。双流融合模块进一步利用历史预测优化动作。在SIMPLER和LIBERO等多个仿真平台及WidowX、Google Robot、Franka等机器人平台上测试表明,FPC-VLA在零样本和微调设置下均优于当前最先进模型。多样化的长时程真实任务部署验证了其出色的泛化能力与实际应用价值,为构建更可靠的自主系统提供了新路径。

原文摘要 · Abstract (English)

Robotic manipulation is a fundamental component of automation. However, traditional perception-planning pipelines often fall short in open-ended tasks due to limited flexibility, while the architecture of a single end-to-end Vision-Language-Action (VLA) offers promising capabilities but lacks crucial mechanisms for anticipating and recovering from failure. To address these challenges, we propose FPC-VLA, a dual-model framework that integrates VLA with a supervisor for failure prediction and correction. The supervisor evaluates action viability through vision-language queries and generates corrective strategies when risks arise, trained efficiently without manual labeling. A dual-stream fusion module further refines actions by leveraging past predictions. Evaluation results on multiple simulation platforms (SIMPLER and LIBERO) and robot embodiments (WidowX, Google Robot, Franka) show that FPC-VLA outperforms state-of-the-art models in both zero-shot and fine-tuned settings. Successful real-world deployments on diverse, long-horizon tasks confirm FPC-VLA's strong generalization and practical utility for building more reliable autonomous systems.

机器人操作视觉语言动作故障预测自主系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。