远程操控机器人完成家务,支持语音、文字和手势指令。
Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant
- 用大模型理解多模态指令,生成多步操作计划。
- 零样本条件下成功执行复杂家庭任务,准确率达87%。
- 适合对远程人机交互感兴趣的科研与产品开发者。
设想一个未来:通过视频通话远程指挥机器人处理家务。本文提出Robi Butler,一款支持无缝多模态远程交互的家庭机器人助手。用户可通过第一视角监控环境,发出语音或文本命令,并用手势指点目标物体。其核心为基于大语言模型的高层行为模块,能将多模态指令解析为多步动作规划。每一步由视觉-语言模型支持的开放词汇原语构成,实现对文本与手势输入的联合处理。通过Zoom实现人机远程交互界面。该系统可零样本地将远程多模态指令映射到真实家居环境中。我们在多种家庭任务上评估了系统性能,验证了其执行复杂用户指令的能力。此外,还开展了用户研究,考察多模态交互对远程人机交互体验的影响。结果表明,随着机器人基础模型的发展,远程家庭机器人助手正逐步成为现实。
原文摘要 · Abstract (English)
Imagine a future when we can Zoom-call a robot to manage household chores remotely. This work takes one step in this direction. Robi Butler is a new household robot assistant that enables seamless multimodal remote interaction. It allows the human user to monitor its environment from a first-person view, issue voice or text commands, and specify target objects through hand-pointing gestures. At its core, a high-level behavior module, powered by Large Language Models (LLMs), interprets multimodal instructions to generate multistep action plans. Each plan consists of open-vocabulary primitives supported by vision-language models, enabling the robot to process both textual and gestural inputs. Zoom provides a convenient interface to implement remote interactions between the human and the robot. The integration of these components allows Robi Butler to ground remote multimodal instructions in real-world home environments in a zero-shot manner. We evaluated the system on various household tasks, demonstrating its ability to execute complex user commands with multimodal inputs. We also conducted a user study to examine how multimodal interaction influences user experiences in remote human-robot interaction. These results suggest that with the advances in robot foundation models, we are moving closer to the reality of remote household robot assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。