大模型能说但不会看时机,难在自然对话中判断何时该接话。
Large Language Models Know What To Say But Not When To Speak
- 构建了人工标注的对话中段发言点数据集,填补研究空白
- 现有大模型在非句尾发言点预测上表现差,准确率不足40%
- 适合研究对话系统、人机交互和语音助手的开发者参考
话语轮换是人类交流中的基础机制,确保口语互动顺畅连贯。近年来大语言模型(LLMs)的发展推动其在语音对话系统(SDS)中提升轮换能力,如适时回应。然而,现有模型普遍难以预测自然非剧本化对话中的发言机会——即过渡相关点(TRPs),仅关注句末TRPs,忽略句内TRPs。为此,我们引入一个由参与者标注的句内TRPs新数据集,并用于评估先进LLMs在预测发言时机上的表现。实验揭示当前大模型在建模非剧本化口语互动方面存在明显局限,指出了改进方向,为更自然的对话系统发展奠定基础。
原文摘要 · Abstract (English)
Turn-taking is a fundamental mechanism in human communication that ensures smooth and coherent verbal interactions. Recent advances in Large Language Models (LLMs) have motivated their use in improving the turn-taking capabilities of Spoken Dialogue Systems (SDS), such as their ability to respond at appropriate times. However, existing models often struggle to predict opportunities for speaking -- called Transition Relevance Places (TRPs) -- in natural, unscripted conversations, focusing only on turn-final TRPs and not within-turn TRPs. To address these limitations, we introduce a novel dataset of participant-labeled within-turn TRPs and use it to evaluate the performance of state-of-the-art LLMs in predicting opportunities for speaking. Our experiments reveal the current limitations of LLMs in modeling unscripted spoken interactions, highlighting areas for improvement and paving the way for more naturalistic dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。