针对中文语音交互,构建了首个细粒度评测基准
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
- 按能力维度拆分,设计10个子集评估中文语音对话能力
- 覆盖12类用户场景,包含超1万条高质量语音实例
- 适合研发语音交互系统的团队与评估模型性能的研究者
多模态大语言模型的发展催生了具备语音交互能力的智能系统。作为全球使用最广泛的语言之一,中文已被多数模型支持以提升其适用性与覆盖范围。然而,当前缺乏针对中文语境的完整语音到语音(S2S)评测基准,阻碍了开发者进行系统化评估,也影响了用户间模型的公平比较。为此,我们提出VocalBench-zh,一个面向中文语境的能力级分解评测套件,包含10个精心设计的子集和超过1万条高质量语音实例,涵盖12类用户导向能力。对14个主流模型的评测实验揭示了现有方法的普遍挑战,并凸显了下一代语音交互系统所需的新洞察。评测代码与数据集将开源于https://github.com/SJTU-OmniAgent/VocalBench-zh。
原文摘要 · Abstract (English)
The development of multi-modal large language models (LLMs) leads to intelligent approaches capable of speech interactions. As one of the most widely spoken languages globally, Mandarin is supported by most models to enhance their applicability and reach. However, the scarcity of comprehensive speech-to-speech (S2S) benchmarks in Mandarin contexts impedes systematic evaluation for developers and hinders fair model comparison for users. In this work, we propose VocalBench-zh, an ability-level divided evaluation suite adapted to Mandarin context consisting of 10 well-crafted subsets and over 10K high-quality instances, covering 12 user-oriented characters. The evaluation experiment on 14 mainstream models reveals the common challenges for current routes, and highlights the need for new insights into next-generation speech interactive systems. The evaluation codes and datasets will be available at https://github.com/SJTU-OmniAgent/VocalBench-zh.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。