首个系统评估语音大模型对话能力的基准测试
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
- 构建涵盖语义与声学生成能力的综合评测框架
- 首次量化评估语音大模型的自然流畅度与音质表现
- 适合语音助手研发者与评估研究人员参考
得益于大语言模型(LLMs)、语音编码算法和声码器结构的持续进步,近期技术已实现直接根据用户指令生成语音响应。然而,生成语音质量的评估长期被忽视,尤其是在从追求语义准确性转向强调自然生动的语音流时。以往评估主要关注语音理解能力,缺乏对声学质量的量化。本文提出语音对话助手基准测试(SOVA-Bench),系统比较现有语音大模型在通用知识、语音识别与理解,以及语义和声学生成能力方面的表现。据我们所知,SOVA-Bench 是目前最系统的语音大模型评估框架之一,为语音交互系统的发展提供了重要方向。
原文摘要 · Abstract (English)
Thanks to the steady progress of large language models (LLMs), speech encoding algorithms and vocoder structure, recent advancements have enabled generating speech response directly from a user instruction. However, benchmarking the generated speech quality has been a neglected but critical issue, considering the shift from the pursuit of semantic accuracy to vivid and spontaneous speech flow. Previous evaluation focused on the speech-understanding ability, lacking a quantification of acoustic quality. In this paper, we propose Speech cOnversational Voice Assistant Benchmark (SOVA-Bench), providing a comprehension comparison of the general knowledge, speech recognition and understanding, along with both semantic and acoustic generative ability between available speech LLMs. To the best of our knowledge, SOVA-Bench is one of the most systematic evaluation frameworks for speech LLMs, inspiring the direction of voice interaction systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。