arXiv:2409.03966cs.RO2024-09被引 23

优化提示词提升视觉语言模型空间推理能力,自动修复机器人故障

Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts

  • 通过优化视觉与文本提示,增强VLM的空间推理能力
  • 运动级纠错准确率提升65.78%,任务级恢复成功率提高5.8%~7.5%
  • 无需人工干预,适用于未知故障的自动化恢复,适合机器人研发者

当前机器人自主性难以突破预设运行设计域(ODD),而现实世界充满不确定性导致故障频发。传统方法依赖人工干预或穷举故障场景设计恢复策略,成本高昂。基础视觉-语言模型(VLM)具备强泛化与推理能力,但空间推理仍受限。本文研究如何通过优化视觉与文本提示,提升VLM在机器人控制中的空间推理能力,使其作为黑箱控制器实现运动级位置修正与任务级未知故障恢复。优化包括识别视觉提示中的关键元素、在文本提示中高亮查询、分解故障检测与控制生成的推理过程。实验表明,优化提示使运动级位置错误纠正准确率显著优于预训练视觉-语言-动作模型,提升65.78%;在乐高积木组装任务中,对未知故障的检测、分析与恢复计划生成成功率分别提升5.8%、5.8%和7.5%。

原文摘要 · Abstract (English)

Current robot autonomy struggles to operate beyond the assumed Operational Design Domain (ODD), the specific set of conditions and environments in which the system is designed to function, while the real-world is rife with uncertainties that may lead to failures. Automating recovery remains a significant challenge. Traditional methods often rely on human intervention to manually address failures or require exhaustive enumeration of failure cases and the design of specific recovery policies for each scenario, both of which are labor-intensive. Foundational Vision-Language Models (VLMs), which demonstrate remarkable common-sense generalization and reasoning capabilities, have broader, potentially unbounded ODDs. However, limitations in spatial reasoning continue to be a common challenge for many VLMs when applied to robot control and motion-level error recovery. In this paper, we investigate how optimizing visual and text prompts can enhance the spatial reasoning of VLMs, enabling them to function effectively as black-box controllers for both motion-level position correction and task-level recovery from unknown failures. Specifically, the optimizations include identifying key visual elements in visual prompts, highlighting these elements in text prompts for querying, and decomposing the reasoning process for failure detection and control generation. In experiments, prompt optimizations significantly outperform pre-trained Vision-Language-Action Models in correcting motion-level position errors and improve accuracy by 65.78% compared to VLMs with unoptimized prompts. Additionally, for task-level failures, optimized prompts enhanced the success rate by 5.8%, 5.8%, and 7.5% in VLMs' abilities to detect failures, analyze issues, and generate recovery plans, respectively, across a wide range of unknown errors in Lego assembly.

机器人视觉语言模型故障恢复提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。