构建首个评估语音对话模型推理、口语化与副语言能力的综合性基准。
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
- 分三部分评测:推理难度、口语自然度、声学特征理解与生成。
- 现有模型在口语流畅性与副语言表现上仍显著落后于人类水平。
- 适合研究语音对话系统、人机交互与多模态智能的学者使用。
随着先进推理能力融入语音对话模型,领域亟需超越简单交互的评估基准以应对真实世界复杂性。然而,现有评价体系主要沿用文本生成标准,忽视了语音对话中副语言特征和口语化表达的独特性,以及现代智能体所需的认知深度。为此,我们提出 WavBench,一个全面评估真实对话能力的基准。其独特之处在于构建三重框架:1)Pro 子集,通过显著提升难度,严格测试推理增强型模型;2)Basic 子集,定义新型口语化标准,强调“可听性”——即自然词汇、语言流畅性与互动亲和力,而非僵化的书面准确性;3)Acoustic 子集,涵盖显性理解、生成及隐含对话,严格评估真实场景下的综合副语言能力。通过对五个前沿模型的评估,WavBench 揭示了复杂问题解决、口语化表达与副语言保真度之间的交叉挑战,为构建稳健的语音对话系统提供关键指引。数据集与评估工具已公开于 https://naruto-2024.github.io/wavbench.github.io/。
原文摘要 · Abstract (English)
With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly adhere to text-generation standards, overlooking the unique audio-centric characteristics of paralinguistics and colloquialisms, alongside the cognitive depth required by modern agents. To bridge this gap, we introduce WavBench, a comprehensive benchmark designed to evaluate realistic conversational abilities where prior works fall short. Uniquely, WavBench establishes a tripartite framework: 1) Pro subset, designed to rigorously challenge reasoning-enhanced models with significantly increased difficulty; 2) Basic subset, defining a novel standard for spoken colloquialism that prioritizes "listenability" through natural vocabulary, linguistic fluency, and interactive rapport, rather than rigid written accuracy; and 3) Acoustic subset, covering explicit understanding, generation, and implicit dialogue to rigorously evaluate comprehensive paralinguistic capabilities within authentic real-world scenarios. Through evaluating five state-of-the-art models, WavBench offers critical insights into the intersection of complex problem-solving, colloquial delivery, and paralinguistic fidelity, guiding the evolution of robust spoken dialogue models. The benchmark dataset and evaluation toolkit are available at https://naruto-2024.github.io/wavbench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。