让机器人通过手势和语音与人协作取货架上的箱子,更安全高效。
A Multi-Modal Interaction Framework for Efficient Human-Robot Collaborative Shelf Picking
- 融合手势、语音与物理仿真,实现多模态交互与智能决策。
- 在真实场景中完成三类协作任务,提升取物效率与稳定性。
- 适合需要人机协同的仓储、物流场景,尤其看重安全性与自然交互。
服务机器人在仓库等以人类为中心的环境中日益普及,亟需实现无缝且直观的人机协作。本文提出一种协同货架取物框架,结合多模态交互、基于物理的推理与任务分工,提升人机团队协作能力。该框架使机器人能识别手指指向动作、理解语音指令,并通过视觉与听觉反馈进行沟通。系统由大型语言模型(LLM)驱动,利用思维链(CoT)与物理仿真引擎,安全地从杂乱堆叠的货架中取出箱子;通过关系图生成子任务、规划提取顺序并做出决策。我们通过三个真实场景实验验证:1)手势引导取箱;2)协同清空货架;3)协同稳定协助。结果表明该框架有效提升了协作效率与安全性。
原文摘要 · Abstract (English)
The growing presence of service robots in human-centric environments, such as warehouses, demands seamless and intuitive human-robot collaboration. In this paper, we propose a collaborative shelf-picking framework that combines multimodal interaction, physics-based reasoning, and task division for enhanced human-robot teamwork. The framework enables the robot to recognize human pointing gestures, interpret verbal cues and voice commands, and communicate through visual and auditory feedback. Moreover, it is powered by a Large Language Model (LLM) which utilizes Chain of Thought (CoT) and a physics-based simulation engine for safely retrieving cluttered stacks of boxes on shelves, relationship graph for sub-task generation, extraction sequence planning and decision making. Furthermore, we validate the framework through real-world shelf picking experiments such as 1) Gesture-Guided Box Extraction, 2) Collaborative Shelf Clearing and 3) Collaborative Stability Assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。