用视觉语言模型统一规划与执行跨建筑的复杂操作任务
BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation
- 基于视觉语言模型整合感知、技能与记忆,实现长时序任务规划
- 在70次不同场景测试中达成47.1%成功率,单次任务最长15分钟
- 适合需要跨楼层、多物体交互的实用型服务机器人研发
为在建筑尺度上运行,服务机器人需完成涉及多房间、多楼层及大量未见过日常物品的长期移动操作任务。我们称此类任务为「跨建筑移动操作」。为应对这类长时序任务,本文提出BUMBLE,一种基于视觉语言模型(VLM)的统一框架,集成开放世界RGBD感知、从粗到细的多种运动技能以及双层记忆机制。大规模评估(90+小时)显示,BUMBLE在需组合最多12个真实技能、每轮任务持续约15分钟的长时序任务中优于多个基线方法。在不同建筑、任务和场景布局下,70次试验平均成功率达47.1%,且起始位置可在不同房间与楼层。用户研究显示,本方法满意度比现有先进方法高出22%。最后,我们展示了利用日益强大的基础模型可进一步提升性能的潜力。
原文摘要 · Abstract (English)
To operate at a building scale, service robots must perform very long-horizon mobile manipulation tasks by navigating to different rooms, accessing different floors, and interacting with a wide and unseen range of everyday objects. We refer to these tasks as Building-wide Mobile Manipulation. To tackle these inherently long-horizon tasks, we introduce BUMBLE, a unified Vision-Language Model (VLM)-based framework integrating open-world RGBD perception, a wide spectrum of gross-to-fine motor skills, and dual-layered memory. Our extensive evaluation (90+ hours) indicates that BUMBLE outperforms multiple baselines in long-horizon building-wide tasks that require sequencing up to 12 ground truth skills spanning 15 minutes per trial. BUMBLE achieves 47.1% success rate averaged over 70 trials in different buildings, tasks, and scene layouts from different starting rooms and floors. Our user study demonstrates 22% higher satisfaction with our method than state-of-the-art mobile manipulation methods. Finally, we demonstrate the potential of using increasingly-capable foundation models to push performance further. For more information, see https://robin-lab.cs.utexas.edu/BUMBLE/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。