arXiv:2603.27667cs.SDcs.AI2026-03

提出证据优先的音频理解框架,提升复杂场景下听觉信息保留能力。

EvA: An Evidence-First Audio Understanding Paradigm for LALMs

论文配图:EvA: An Evidence-First Audio Understanding Paradigm for LALMs
图 1 · 摘自论文原文
  • 采用双路径架构,通过分层聚合与时间对齐融合强化声学证据保存
  • 在MMAU、MMAR、MMSU上零样本测试中取得最佳开源感知性能,感知密集任务提升显著
  • 适用于需要精细声学理解的任务,如开放问答与细粒度音频描述生成

大型音频语言模型在复杂声学场景中仍表现不佳,因其常在推理前丢失任务相关的声学证据。我们识别出这一问题为‘证据瓶颈’:当前先进系统在声学证据提取上的缺陷大于下游推理能力不足,表明上游感知是主要限制因素。为此,我们提出EvA(Evidence-First Audio)——一种双路径架构,通过分层聚合与非压缩、时间对齐融合增强声学证据保留。同时构建了EvA-Perception,一个包含约54,000条事件顺序描述和500,000个基于证据的问答对的大规模训练数据集。在统一零样本协议下,EvA在MMAU、MMAR、MMSU上达到最优开源感知表现,尤其在感知密集子集上提升最大。人类评估显示其开放性描述任务中具备更细粒度的声学覆盖与更高质量。结果支持‘证据优先’假设:更强的音频理解依赖于推理前的声学证据保持。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We identify this error pattern as the evidence bottleneck: state-of-the-art systems show larger deficits in acoustic evidence extraction than in downstream reasoning, suggesting that upstream perception is often the limiting factor. To address this problem, we propose EvA (Evidence-First Audio), a dual-path architecture that enhances acoustic evidence preservation through hierarchical aggregation and non-compressive, time-aligned fusion. We also build EvA-Perception, a large-scale training set with about 54K event-ordered captions and 500K evidence-grounded QA pairs. Under a unified zero-shot protocol, EvA achieves the best open-source \emph{Perception} results on MMAU, MMAR, and MMSU, with the largest gains on perception-heavy splits. Human evaluation on open-ended captioning further shows improved fine-grained acoustic coverage and caption quality. These results support the evidence-first hypothesis: stronger audio understanding depends on preserving acoustic evidence before reasoning. Project can be found at https://satsuki2486441738.github.io/EvA/.

音频理解证据优先双路径零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。