用对话模型让机器人更自然地轮流说话,减少停顿和打断。
Applying General Turn-taking Models to Conversational Human-Robot Interaction
- 引入通用对话轮换模型,无需微调即可预测发言时机。
- 实验显示响应延迟减少,用户对轮换体验的满意度显著提升。
- 适合研究人机对话交互、希望提升对话流畅性的开发者。
对话中的轮换是交流的基础,但当前人机交互系统多依赖简单的静默判断模型,导致不自然的停顿与打断。本文首次探究通用轮换模型(TurnGPT 和语音活动投影,VAP)在人机对话中的应用。这些模型基于人类对话数据,通过自监督学习训练,无需领域微调。我们提出将二者协同使用的方法,以预测机器人何时应准备回应、切入对话或处理打断。在包含39名成人的受控实验中,使用Furhat机器人结合大语言模型进行自主回复生成,结果表明参与者更偏好新系统,且响应延迟和打断次数均显著降低。
原文摘要 · Abstract (English)
Turn-taking is a fundamental aspect of conversation, but current Human-Robot Interaction (HRI) systems often rely on simplistic, silence-based models, leading to unnatural pauses and interruptions. This paper investigates, for the first time, the application of general turn-taking models, specifically TurnGPT and Voice Activity Projection (VAP), to improve conversational dynamics in HRI. These models are trained on human-human dialogue data using self-supervised learning objectives, without requiring domain-specific fine-tuning. We propose methods for using these models in tandem to predict when a robot should begin preparing responses, take turns, and handle potential interruptions. We evaluated the proposed system in a within-subject study against a traditional baseline system, using the Furhat robot with 39 adults in a conversational setting, in combination with a large language model for autonomous response generation. The results show that participants significantly prefer the proposed system, and it significantly reduces response delays and interruptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。