剖析大模型在音频时序理解中的失败原因,发现注意力分配比模态不平衡更关键。
A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

- 构建1657个问题的基准,专用于分析音频时序推理机制。
- 调整音频注意力分布比单纯放大注意力更有效,提升准确率至59.1%。
- 适合研究多模态模型失效机制与注意力优化的研究者。
大型音频语言模型(LALMs)在多种音频理解任务中表现强劲,但在时序推理方面仍存在显著缺陷,而这一能力是人类听觉感知的核心。现有基准仅报告性能差距,未深入探究根本原因。为此,我们设计了一个包含1,657个问题的基准,覆盖三个基础任务,专为机制分析而设。通过行为分析发现,当文本线索存在时,模型常过度依赖文本而忽视音频。我们首次对LALMs的时序推理失败进行因果机制分析:对比注意力上权重与扩大规模,发现重新分配音频令牌注意力比增加注意力更有效;聚焦任务相关令牌可进一步提升性能。这些结果表明,单一模态失衡无法解释失败。在瓶颈层进行注意力缩放,无需微调即可将准确率从55.9%提升至59.1%,为未来工作指明方向。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception. Understanding the causes of these failures remains challenging as existing benchmarks report performance gaps without probing underlying mechanisms. To address this, we introduce a benchmark with 1,657 questions across three foundational tasks designed specifically for mechanistic analysis. Examining model outputs across varying input settings (behavioral analysis) reveals that models often under-utilize audio when textual cues are available. We also provide the first causal mechanistic analysis of temporal reasoning failures in LALMs. Comparing attention upweighting against scaling, we find that redistributing attention across audio tokens is more effective than increasing audio attention. Targeting task-relevant tokens yields further gains. These findings suggest that modality imbalance alone cannot explain failures. Attention scaling at bottleneck layers improves accuracy from 55.9% to 59.1% without fine-tuning, demonstrating a promising direction for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。