构建多场景对话轮换评估基准,揭示模型在不同情境下的表现差异
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
- 将轮换决策建模为结构化问题,设计细粒度分类与可控上下文变化
- 在统一协议下测试主流模型,发现决策类型与交互场景间性能差异显著
- 适合研究对话系统评估、人机交互与自然语言理解的学者使用
对话轮换建模是语音对话系统的核心,但现有评估方法零散且局限于狭窄交互场景下的二元边界检测,难以进行系统比较并掩盖模型在不同对话条件下的弱点。我们提出CoDeTT,一个面向轮换决策的上下文感知评估基准。CoDeTT将轮换问题形式化为结构化决策任务,构建了一个包含多场景、细粒度决策类别与受控上下文变化的数据集。在统一评估协议下,我们对代表性模型进行了评估,发现不同决策类型和交互场景间存在显著性能差异。CoDeTT为轮换系统的系统性、上下文感知评估提供了标准化基准。数据集与评估工具已开源:https://yingaowang-casia.github.io/CoDeTT.github.io/
原文摘要 · Abstract (English)
Turn-taking modeling is fundamental to spoken dialogue systems, yet its evaluation remains fragmented and often limited to binary boundary detection under narrow interaction settings. Such protocols hinder systematic comparison and obscure model weaknesses across conversational conditions. We present CoDeTT, a context-aware decision benchmark for turn-taking evaluation. CoDeTT formulates turn-taking as a structured decision problem and constructs a multi-scenario dataset with fine-grained decision categories and controlled context variations. Under a unified evaluation protocol, we assess representative existing models and observe substantial performance disparities across decision types and interaction scenarios. CoDeTT provides a standardized benchmark for systematic and context-aware evaluation of turn-taking systems. The benchmark dataset and evaluation toolkit are available at https://yingaowang-casia.github.io/CoDeTT.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。