首个基于真实中文语音的语音对话模型评测基准,全面评估模型能力。
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
- 构建全真人语音的中文评测数据集,覆盖指令遵循、知识理解与鲁棒性三维度。
- 在多个真实场景下发现主流模型存在显著性能差距,尤其在语音控制和环境扰动中表现弱。
- 适合语音助手、多模态对话系统研发者使用,推动中文语音交互技术进步。
近年来,大音频语言模型(LALMs)显著提升了多模态对话系统的性能。然而,现有评测基准仍存在局限:以英语为主、依赖合成语音,且缺乏多维度的精细评估。为此,我们提出语音聊天机器人评测基准(VCB Bench)——一个基于真实人类语音的高质量中文评测基准。VCB Bench 从三个互补视角评估 LALMs:指令遵循(包括超越文本指令的语音级控制)、知识理解(通用知识、推理与日常对话)以及鲁棒性(在内容、环境与说话人特征扰动下的稳定性)。对代表性 LALMs 的实验揭示了显著的性能差距,并指明未来改进方向。VCB Bench 提供可复现、细粒度的评估框架,提供标准化方法与实用洞察,助力中文语音对话模型发展。
原文摘要 · Abstract (English)
Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks remain limited -- they are mainly English-centric, rely on synthetic speech, and lack comprehensive, discriminative evaluation across multiple dimensions. To address these gaps, we present Voice Chat Bot Bench (VCB Bench) -- a high-quality Chinese benchmark built entirely on real human speech. VCB Bench evaluates LALMs from three complementary perspectives: instruction following (including speech-level control beyond text commands), knowledge understanding (general knowledge, reasoning, and daily dialogue), and robustness (stability under perturbations in content, environment, and speaker traits). Experiments on representative LALMs reveal notable performance gaps and highlight future directions for improvement. VCB Bench provides a reproducible and fine-grained evaluation framework, offering standardized methodology and practical insights for advancing Chinese voice conversational models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。