arXiv:2505.03054cs.AIcs.CL2025-05被引 10

评测大模型对长达51分钟语音的理解能力,发现现有模型普遍表现不佳。

BLAB: Brutally Long Audio Bench

  • 构建时长平均51分钟的长音频评测集,涵盖定位、计数等四项任务。
  • 6个主流语音大模型在长音频上表现均差,性能随时长增加显著下降。
  • 适合关注长时序语音理解的开发者和研究者,推动更鲁棒模型设计。

构建大型语音语言模型(LM)以理解多样化的口语交互,对于适应人类沟通的多模态特性至关重要,并能提升语言技术对不同用户群体的可及性。当前语音LM的研究主要集中在30秒以内的短音频片段,对更贴近真实交互的长篇对话音频探索不足。我们提出 Brutally Long Audio Bench(BLAB),一个挑战性的长音频基准测试,评估语音LM在定位、时长估计、情绪识别和计数任务上的表现,使用平均时长为51分钟的音频片段。BLAB包含超过833小时的多样化全时长音频,每段音频均配有由人工标注的自然语言问答对。音频数据来源于许可宽松的来源,并经过人工筛选以确保任务合规性。我们在六种开源与专有语音LM上评估了BLAB,发现包括Gemini 2.0 Pro和GPT-4o在内的先进模型均表现不佳。综合分析揭示任务难度与音频时长之间的权衡:语音LM在长音频上普遍表现差,性能随时长增加而下降;在定位、时间推理、计数任务上表现弱,且依赖提示而非音频内容理解非语音信息。BLAB为发展具备强大长音频理解能力的语音模型提供了挑战性评估框架。

原文摘要 · Abstract (English)

Developing large audio language models (LMs) capable of understanding diverse spoken interactions is essential for accommodating the multimodal nature of human communication and can increase the accessibility of language technologies across different user populations. Recent work on audio LMs has primarily evaluated their performance on short audio segments, typically under 30 seconds, with limited exploration of long-form conversational speech segments that more closely reflect natural user interactions with these models. We introduce Brutally Long Audio Bench (BLAB), a challenging long-form audio benchmark that evaluates audio LMs on localization, duration estimation, emotion, and counting tasks using audio segments averaging 51 minutes in length. BLAB consists of 833+ hours of diverse, full-length audio clips, each paired with human-annotated, text-based natural language questions and answers. Our audio data were collected from permissively licensed sources and underwent a human-assisted filtering process to ensure task compliance. We evaluate six open-source and proprietary audio LMs on BLAB and find that all of them, including advanced models such as Gemini 2.0 Pro and GPT-4o, struggle with the tasks in BLAB. Our comprehensive analysis reveals key insights into the trade-offs between task difficulty and audio duration. In general, we find that audio LMs struggle with long-form speech, with performance declining as duration increases. They perform poorly on localization, temporal reasoning, counting, and struggle to understand non-phonemic information, relying more on prompts than audio content. BLAB serves as a challenging evaluation framework to develop audio LMs with robust long-form audio understanding capabilities.

语音理解长音频大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。