arXiv:2503.15496cs.HCcs.RO2025-03被引 4

让机器人在多人对话中准确识别说话人并自然回应。

Fast Multi-Party Open-Ended Conversation with a Social Robot

  • 融合语音方向、说话人辨识和人脸识别,用大模型生成回应。
  • 并行对话中指认准确率达92.6%,人脸辨识可靠度80%-94%。
  • 适合研究社交机器人交互与多模态对话系统的人参考。

多人群体开放式对话仍是人机交互中的重大挑战,尤其在机器人需识别说话人、分配发言权并应对重叠或快速切换对话时。本文提出一种结合多模态感知(声源定位、说话人辨识、人脸辨识)与大语言模型响应生成的多人群体对话系统。该系统部署于Furhat机器人,在30名参与者中通过两种场景评估:(i) 并行独立对话,(ii) 共享群体讨论。结果表明,系统能保持连贯且吸引人的对话,在并行场景中实现92.6%的高指认准确率,人脸辨识可靠性达80%-94%。参与者报告了清晰的社会存在感和积极互动体验,但基于音频的说话人识别错误与响应延迟仍影响群体互动流畅性。结果凸显了基于大模型的多人群体交互的潜力与局限,并为未来社交机器人在多模态线索融合与响应速度方面的改进指明具体方向。

原文摘要 · Abstract (English)

Multi-party open-ended conversation remains a major challenge in human-robot interaction, particularly when robots must recognise speakers, allocate turns, and respond coherently under overlapping or rapidly shifting dialogue. This paper presents a multi-party conversational system that combines multimodal perception (voice direction of arrival, speaker diarisation, face recognition) with a large language model for response generation. Implemented on the Furhat robot, the system was evaluated with 30 participants across two scenarios: (i) parallel, separate conversations and (ii) shared group discussion. Results show that the system maintains coherent and engaging conversations, achieving high addressee accuracy in parallel settings (92.6%) and strong face recognition reliability (80-94%). Participants reported clear social presence and positive engagement, although technical barriers such as audio-based speaker recognition errors and response latency affected the fluidity of group interactions. The results highlight both the promise and limitations of LLM-based multi-party interaction and outline concrete directions for improving multimodal cue integration and responsiveness in future social robots.

社交机器人多模态对话系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。