首个纯语音评测框架,专为评估语音模型知识理解能力而设计
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
- 全程使用语音输入输出,真实模拟人机语音交互
- 在多种音频条件下测试模型表现,揭示其对噪声敏感性
- 首次评测语音模式下的数学推理能力,推动模型升级
随着语音交互需求上升,端到端语音语言模型(SLMs)成为有前景的解决方案。然而,现有问答基准无法支持纯语音评估,也未考虑多样化的音频输入条件,难以有效衡量模型的知识理解能力。为此,我们提出VoxEval——一个全新的语音问答基准,通过纯语音交互评估SLMs的知识理解。该基准具备三大特性:1)保持输入与输出均为语音格式;2)在多种音频条件下评估模型鲁棒性;3)首次实现语音形式下的复杂任务(如数学推理)评估。系统性实验表明,当前SLMs在VoxEval上面临显著挑战,暴露出对音频变化的敏感性,并凸显未来需增强推理能力。VoxEval数据集已公开于https://github.com/dreamtheater123/VoxEval。
原文摘要 · Abstract (English)
With the rising need for speech-based interaction models, end-to-end Spoken Language Models (SLMs) have emerged as a promising solution. While these models require comprehensive world knowledge for meaningful and reliable human interactions, existing question-answering (QA) benchmarks fall short in evaluating SLMs' knowledge understanding due to their inability to support end-to-end speech evaluation and account for varied input audio conditions. To address these limitations, we present VoxEval, a novel SpeechQA benchmark that assesses SLMs' knowledge understanding through pure speech interactions. Our benchmark 1) uniquely maintains speech format for both inputs and outputs, 2) evaluates model robustness across diverse input audio conditions, and 3) pioneers the assessment of complex tasks like mathematical reasoning in spoken format. Systematic evaluation demonstrates that VoxEval presents significant challenges to current SLMs, revealing their sensitivity to varying audio conditions and highlighting the need to enhance reasoning capabilities in future development. We hope this benchmark could guide the advancement of more sophisticated and reliable SLMs. VoxEval dataset is available at: https://github.com/dreamtheater123/VoxEval
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。