构建多领域对话轮次基准,揭示系统在打断检测上的类型依赖缺陷
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

- 基于30小时人工标注对话数据,设计标准化评估协议
- 发现打断误报率在回音密集场景中显著升高,且与互动类型强相关
- 适合研究语音交互、人机对话系统的开发与评估人员
自然对话中说话人实时决定何时接话、保持或让出话语权。然而,由于缺乏一致的、基于语言学的评估协议和覆盖多种对话类型的标注数据,轮次管理评估仍受限。为此,我们提出TurnBench,一个包含30小时双人对话手标语料和标准化端轮次/打断检测评估协议的多领域基准。以对话类型为可控变量,涵盖六种不同互动风格,并对每段对话进行三重标注。在14种异构轮次管理系统上进行基准测试,发现端轮次召回率在各类别间稳定,而打断误报率强烈依赖于互动类型,集中出现在回音密集型互动中。尽管人类听者在当前话语结束前中位151毫秒就开始发言,但现有系统无法在不产生大量误报的情况下实现同等表现。我们发布该语料库、104小时训练集及公开排行榜,含交互式数据查看器,网址:https://turnbench.sesame.com
原文摘要 · Abstract (English)
Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically grounded evaluation protocol and hand-annotated data covering diverse conversation types. To address this, we present TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection. We set conversation type as a controllable experimental variable, covering six distinct interaction styles, and triple-annotate each conversation. Benchmarking 14 heterogeneous turn-taking systems, we find end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles. Although in smooth floor transfers human listeners begin speaking a median 151 ms before the current turn ends, no current system performs equivalently without incurring excessive false positives. We release our corpus, a 104-hour training set, and a public leaderboard with an interactive dataset viewer at https://turnbench.sesame.com
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。