arXiv:2605.27039eess.AScs.SD2026-05被引 2

揭示大模型听觉记忆失效主因:表征漂移而非注意力分配

Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory

论文配图:Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory
图 1 · 摘自论文原文
  • 构建可控多轮听觉任务基准EnvMem,分离语义与声学信息处理
  • 发现非语言信息在表征空间中持续漂移是性能下降主因
  • 适合研究长上下文音频模型、听觉记忆机制的学者参考

大型音频语言模型(LALMs)能处理语音和环境声学线索,但在多轮交互中难以保留非语音信息。语义(语音)与声学(非语音)理解之间的性能差距尚不明确,其表征与检索机制仍不清楚。本文提出EnvMem,一个受控的多轮基准,用于研究这一差距并识别表征(潜在嵌入)与检索(注意力分配)层面失败的根本原因。通过事后干预,我们探测了表征结构与注意力动态。结果表明,表征轨迹漂移是主要失败模式,而注意力分配对性能下降的解释作用有限。整体上,本文提供了系统分析与改进长上下文LALMs非语言记忆的框架,为未来数据与训练设计提供启示。

原文摘要 · Abstract (English)

Large audio language models (LALMs) process both speech and environmental acoustic cues, yet struggle to retain non-speech information across multi-turn interactions. The performance gap between semantic (speech) and acoustic (non-speech) understanding remains poorly understood, and the underlying mechanisms of representation and retrieval are still unclear. This work introduces EnvMem, a controlled multi-turn benchmark designed to study this gap and identify the root causes of failures at the representation (i.e., latent embeddings) and retrieval levels (i.e., attention allocation). We further conduct post-hoc interventions to probe representational structure and attention dynamics. Our results reveal representational trajectory drift as the key failure mode, while showing that attention allocation plays a limited role in explaining the observed degradation. Overall, we provide a systematic framework for analyzing and improving non-linguistic memory in long-context LALMs, shedding light on future data and training design for robust acoustic memory modeling.

音频模型听觉记忆表征漂移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。