首个综合评估语音、场景与事件理解能力的音频大模型基准
Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
- 构建考虑语音与非语音能量差异的多任务评测框架
- 发现多数模型在联合理解任务中表现下降,尤其在复杂场景下
- 提出思维链方法,显著提升跨模态联合推理能力
近年来,大型音频语言模型(LALMs)进展迅速,在跨模态融合基础上展现出强大的通用音频理解能力。然而现有基准未能充分覆盖真实场景中的关键问题:音频常同时包含语音与非语音成分,且两者能量水平在不同情境下差异显著;此外,多数评测未涵盖同一音频片段中语音、场景与事件的联合理解。为此,本文提出SSEU-Bench,首个兼顾语音、场景与事件独立与联合理解的多功能音频理解基准,明确考虑语音与非语音成分的能量差异。实验表明,部分LALMs在联合理解设置下性能明显下降。为解决此问题,我们引入思维链(Chain-of-Thought)策略,通过将复杂任务分解为逐步推理步骤,有效提升模型在联合理解任务中的表现。
原文摘要 · Abstract (English)
Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To evaluate LALMs' audio understanding performance, researchers have proposed different benchmarks. However, key aspects for real-world interactions are underexplored in existing benchmarks, i.e., audio signals typically contain both speech and non-speech components, and energy levels of these components can vary significantly across different scenarios. Moreover, most benchmarks do not consider the joint understanding of speech, scene, and events within the same audio clip. In this work, we introduce SSEU-Bench, the first versatile audio understanding benchmark that explicitly accounts for energy differences between speech and non-speech audio, with both independent and joint understanding settings for speech, scene, and events. Furthermore, we demonstrate that some LALMs tend to underperform on certain tasks in a joint understanding setting. To address this issue, we introduce Chain-of-Thought, which effectively improves LALMs' joint audio understanding performance by decomposing complex tasks into simpler reasoning steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。