arXiv:2509.22651cs.CLcs.AI2025-09被引 9

首个全面评估语音助手听、说、看能力的基准测试。

VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing

  • 构建涵盖13类任务的10497个样本评测集。
  • 中等规模模型在听觉理解上超越大模型两倍以上。
  • 适合研究语音交互与多模态系统的学生和开发者。

大型语言模型和多模态系统的快速发展推动了以语音为核心的AI助手兴起,但现有评测基准无法全面评估其能力。我们提出VoiceAssistant-Eval,一个覆盖听、说、看三方面的综合性评测基准,包含10,497个精心筛选的示例,涵盖13个任务类别:听觉方面包括自然声音、音乐和口语对话;说话方面包括多轮对话、角色扮演模仿及多种场景;视觉方面则涉及高度异构的图像。为验证其有效性,我们评估了21个开源模型和GPT-4o-Audio,衡量响应内容质量、语音表现及一致性。结果揭示三个关键发现:(1)专有模型并非普遍优于开源模型;(2)多数模型在说话任务表现优异,但在音频理解上落后;(3)设计得当的小模型可媲美大模型。值得注意的是,中等规模的Step-Audio-2-mini(7B)在听觉理解准确率上超过LLaMA-Omni2-32B-Bilingual的两倍。然而挑战依然存在:多模态输入与角色扮演语音模仿任务对当前模型仍具难度,鲁棒性与安全对齐差距显著。VoiceAssistant-Eval识别出这些短板,并建立严谨框架,指导下一代语音助手的发展。代码与数据将在https://mathllm.github.io/VoiceAssistantEval/ 公开。

原文摘要 · Abstract (English)

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce VoiceAssistant-Eval, a comprehensive benchmark designed to assess AI assistants across listening, speaking, and viewing. VoiceAssistant-Eval comprises 10,497 curated examples spanning 13 task categories. These tasks include natural sounds, music, and spoken dialogue for listening; multi-turn dialogue, role-play imitation, and various scenarios for speaking; and highly heterogeneous images for viewing. To demonstrate its utility, we evaluate 21 open-source models and GPT-4o-Audio, measuring the quality of the response content and speech, as well as their consistency. The results reveal three key findings: (1) proprietary models do not universally outperform open-source models; (2) most models excel at speaking tasks but lag in audio understanding; and (3) well-designed smaller models can rival much larger ones. Notably, the mid-sized Step-Audio-2-mini (7B) achieves more than double the listening accuracy of LLaMA-Omni2-32B-Bilingual. However, challenges remain: multimodal (audio plus visual) input and role-play voice imitation tasks are difficult for current models, and significant gaps persist in robustness and safety alignment. VoiceAssistant-Eval identifies these gaps and establishes a rigorous framework for evaluating and guiding the development of next-generation AI assistants. Code and data will be released at https://mathllm.github.io/VoiceAssistantEval/ .

语音助手多模态评测基准测试音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。