评测音频大模型的对话轮次衔接能力,发现现有系统常抢话或沉默。
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
- 用人工标注的判别模型评估语音模型的轮次切换表现
- 多款模型在判断何时说话上表现不佳,常中断或长时间沉默
- 适合关注语音对话系统交互流畅性的研究者和开发者
近期涌现的音频基础模型(Audio Foundation Models, FMs)为对话建模带来了新可能。然而,目前缺乏对这些模型在自然、互动对话中轮次衔接能力的全面评估。要实现与用户的有意义对话,我们期望模型能流畅地进行发言交替,避免过度重叠或过长沉默。为此,我们提出一种新颖的评估协议,利用一个经过人类-人类对话轮次事件训练的监督判别模型作为评判标准,评估语音对话系统的轮次衔接能力。基于该协议,我们开展了首个综合性用户研究,评估现有语音对话系统在执行轮次切换事件上的表现,揭示出许多有趣现象:模型有时无法理解何时该发言,会过于激进地打断,且很少使用回应性语气(backchannel)。我们进一步在从Switchboard等精心构建的测试基准上,评估多个开源及专有音频基础模型通过API访问的表现,发现其在理解与预测轮次切换方面仍有显著提升空间。我们将开源评估平台,以推动先进对话人工智能系统的发展。
原文摘要 · Abstract (English)
The recent wave of audio foundation models (FMs) could provide new capabilities for conversational modeling. However, there have been limited efforts to evaluate these audio FMs comprehensively on their ability to have natural and interactive conversations. To engage in meaningful conversation with the end user, we would want the FMs to additionally perform a fluent succession of turns without too much overlapping speech or long stretches of silence. Inspired by this, we ask whether the recently proposed audio FMs can understand, predict, and perform turn-taking events? To answer this, we propose a novel evaluation protocol that can assess spoken dialog system's turn-taking capabilities using a supervised model as a judge that has been trained to predict turn-taking events in human-human conversations. Using this protocol, we present the first comprehensive user study that evaluates existing spoken dialogue systems on their ability to perform turn-taking events and reveal many interesting insights, such as they sometimes do not understand when to speak up, can interrupt too aggressively and rarely backchannel. We further evaluate multiple open-source and proprietary audio FMs accessible through APIs on carefully curated test benchmarks from Switchboard to measure their ability to understand and predict turn-taking events and identify significant room for improvement. We will open source our evaluation platform to promote the development of advanced conversational AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。