arXiv:2506.21875cs.CL2025-06被引 10

首个面向真实场景的语音大模型评测基准,助力提升语音交互体验。

WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild

  • 构建真实对话数据集,覆盖多样说话人与复杂声学环境。
  • 引入语音特有现象(如口吃、同音词)测试模型鲁棒性。
  • 设计查询感知评估法,实现细粒度性能分析,适合研发人员使用。

近年来,GPT-4o等多模态大语言模型展现出直接语音交互的强大能力。然而,缺乏专门且全面的端到端语音大模型评测基准,制约了语音大模型在真实应用中的用户体验优化。现有评估方法多沿用文本基准,忽视了语音特有的挑战,如语调、同音词、口吃及用户期望差异。为此,我们首次提出一个系统性评测端到端语音大模型的实际对话能力的综合性基准。通过系统收集与口语场景相关的实际聊天数据,引入说话人属性与声学条件的多样性,并增强语音特有现象。进一步设计查询感知评估方法,采用定制化评估清单与提示,提升自动评估准确性。对主流语音模型进行全面测试与深入分析,揭示模型在不同语音场景下表现差异显著。查询感知评估使细粒度分析成为可能,为语音模型研发与评估提供重要参考。

原文摘要 · Abstract (English)

Recent multi-modal Large Language Models (LLMs) such as GPT-4o have demonstrated strong capabilities of direct speech interaction. However, the lack of specialized and comprehensive benchmarks for end-to-end speech LLM evaluation hinders optimizing the user experience of Audio LLMs in real-world applications. Existing evaluation methods often adapt text-based benchmarks, overlooking speech's unique characteristics and challenges, including prosody, homophones, stuttering, and differing user expectations. Here, we introduce the first comprehensive benchmark designed to systematically evaluate end-to-end speechLLMs in practical speech conversations. We systematically curate real-world chat data relevant to spoken scenarios, introduce diversity in speaker attributes and acoustic conditions, and augment the dataset with speech-specific phenomena. We further design a query-aware evaluation method to use customized evaluation checklists and prompts to enhance the accuracy of automatic evaluation. We conduct comprehensive testing and detailed analysis of various mainstream speech models, revealing significant differences in model performance across different speech scenarios. The use of query-aware evaluation further enables a finer-grained assessment under various speech-specific scenarios. Our benchmark can provide valuable insights for speech model development and evaluation.

语音大模型评测基准端到端真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。