让多个机器人协同与人自然互动,靠感知+语言模型+身体动作统一决策。
A Multimodal Framework for Human-Multi-Agent Interaction
- 每个机器人用多模态感知和大模型实现自主决策,具身化行动。
- 通过中心协调机制避免语音冲突和动作打架,实现流畅交互。
- 适合研究多机器人协作、人机交互的科研人员或开发者参考。
人机交互正向多机器人、社会性场景发展。现有系统难以在统一框架中整合多模态感知、具身表达与协同决策,限制了共享物理空间中的自然与可扩展交互。本文提出一种多模态人-多智能体交互框架,每个机器人作为具备多模态感知和基于大语言模型(LLM)规划的自主认知代理,其决策根植于具身性。团队层面采用中心化协调机制,调控发言顺序与参与度,避免语音重叠和行为冲突。在两台类人机器人上实现,框架通过结合言语、手势、眼神与移动的交互策略,实现连贯的多智能体交互。代表性交互演示展示了跨智能体的协同多模态推理与具身化响应。未来工作将聚焦大规模用户研究及社会性多智能体交互动态的深入探索。
原文摘要 · Abstract (English)
Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a unified framework. This limits natural and scalable interaction in shared physical spaces. We address this gap by introducing a multimodal framework for human-multi-agent interaction in which each robot operates as an autonomous cognitive agent with integrated multimodal perception and Large Language Model (LLM)-driven planning grounded in embodiment. At the team level, a centralized coordination mechanism regulates turn-taking and agent participation to prevent overlapping speech and conflicting actions. Implemented on two humanoid robots, our framework enables coherent multi-agent interaction through interaction policies that combine speech, gesture, gaze, and locomotion. Representative interaction runs demonstrate coordinated multimodal reasoning across agents and grounded embodied responses. Future work will focus on larger-scale user studies and deeper exploration of socially grounded multi-agent interaction dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。