用视觉符号诊断机器人操作失败并提供修复指导
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- 通过显式视觉符号提升故障标注效率
- 构建包含5202条真实轨迹的5.8万组VQA数据集
- 支持闭开两种题型评估,适合机器人自学习研究
视觉-语言-动作(VLA)模型在机器人操作中取得显著进展,但仍受限于故障诊断与从失败中学习的能力。现有失败数据集多为仿真生成,泛化性差。为此,我们提出ViFailback框架,实现机器人操作失败的诊断,并提供文本与视觉修正指引。该框架利用显式视觉符号提升标注效率。我们进一步发布ViFailback数据集,包含58,126组视觉问答对及对应的5,202条真实世界操作轨迹。基于此数据集,构建了ViFailback-Bench基准,涵盖11个细粒度VQA任务,分为面向封闭式问题的ViFailback-Bench Lite与面向开放式问题的ViFailback-Bench Hard。为验证框架有效性,我们构建了ViFailback-8B VLM模型,在ViFailback-Bench上表现显著提升,并能生成用于修正动作的视觉符号。将ViFailback-8B与VLA模型结合后,真实机器人实验表明其具备帮助VLA模型从失败中恢复的能力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic manipulation, yet they remain limited in failure diagnosis and learning from failures. Additionally, existing failure datasets are mostly generated programmatically in simulation, which limits their generalization to the real world. In light of these, we introduce ViFailback, a framework designed to diagnose robotic manipulation failures and provide both textual and visual correction guidance. Our framework utilizes explicit visual symbols to enhance annotation efficiency. We further release the ViFailback dataset, a large-scale collection of 58,126 Visual Question Answering (VQA) pairs along with their corresponding 5,202 real-world manipulation trajectories. Based on the dataset, we establish ViFailback-Bench, a benchmark of 11 fine-grained VQA tasks designed to assess the failure diagnosis and correction abilities of Vision-Language Models (VLMs), featuring ViFailback-Bench Lite for closed-ended and ViFailback-Bench Hard for open-ended evaluation. To demonstrate the effectiveness of our framework, we built the ViFailback-8B VLM, which not only achieves significant overall performance improvement on ViFailback-Bench but also generates visual symbols for corrective action guidance. Finally, by integrating ViFailback-8B with a VLA model, we conduct real-world robotic experiments demonstrating its ability to assist the VLA model in recovering from failures. Project Website: https://x1nyuzhou.github.io/vifailback.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。