构建 VocalBench 基准测试,评估语音交互模型的对话能力
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
- 设计涵盖中英文24000个实例的多维度评估集
- 发现主流27个模型在语义与抗噪能力上普遍不足
- 适合语音交互系统研发与评测人员使用
语音大语言模型(SpeechLLMs)将人机交互从文本拓展至动态语音领域。口语对话包含语义概念、声学变化、副语言线索及环境上下文等多元信息。然而现有评估缺乏真实场景模拟,且仅关注单一性能指标,难以全面比较当前模型的关键能力。为此,我们提出 VocalBench,用于评估语音对话能力,包含约24,000个精心构建的英汉双语实例,覆盖语义质量、声学表现、对话能力与鲁棒性四大维度,涵盖14类用户导向角色。在27个主流模型上的实验揭示了当前方法的共性挑战,凸显下一代语音交互系统需新洞察。
原文摘要 · Abstract (English)
Speech large language models (SpeechLLMs) have extended human-machine interactions from the text modality to the dynamic speech domain. Spoken dialogues convey diverse information, including semantic concepts, acoustic variations, paralanguage cues, and environmental context. However, existing evaluations of speech interaction models lack instances mimicking real scenarios and predominantly focus on the performance of distinct aspects, lacking a comprehensive comparison of critical capabilities between current routines. To address this gap, we propose VocalBench to assess the speech conversational abilities, comprising around 24k carefully curated instances of both English and Mandarin across four key dimensions - semantic quality, acoustic performance, conversational abilities, and robustness, covering 14 user-oriented characters. Experiments on 27 mainstream models reveal the common challenges for current routes, and highlight the need for new insights into next-generation speech interactive systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。