揭示大模型听觉记忆失效主因:表征漂移而非注意力分配
Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory

- 构建可控多轮听觉任务基准EnvMem,分离语义与声学信息处理
- 发现非语言信息在表征空间中持续漂移是性能下降主因
- 适合研究长上下文音频模型、听觉记忆机制的学者参考
大型音频语言模型(LALMs)能处理语音和环境声学线索,但在多轮交互中难以保留非语音信息。语义(语音)与声学(非语音)理解之间的性能差距尚不明确,其表征与检索机制仍不清楚。本文提出EnvMem,一个受控的多轮基准,用于研究这一差距并识别表征(潜在嵌入)与检索(注意力分配)层面失败的根本原因。通过事后干预,我们探测了表征结构与注意力动态。结果表明,表征轨迹漂移是主要失败模式,而注意力分配对性能下降的解释作用有限。整体上,本文提供了系统分析与改进长上下文LALMs非语言记忆的框架,为未来数据与训练设计提供启示。
原文摘要 · Abstract (English)
Large audio language models (LALMs) process both speech and environmental acoustic cues, yet struggle to retain non-speech information across multi-turn interactions. The performance gap between semantic (speech) and acoustic (non-speech) understanding remains poorly understood, and the underlying mechanisms of representation and retrieval are still unclear. This work introduces EnvMem, a controlled multi-turn benchmark designed to study this gap and identify the root causes of failures at the representation (i.e., latent embeddings) and retrieval levels (i.e., attention allocation). We further conduct post-hoc interventions to probe representational structure and attention dynamics. Our results reveal representational trajectory drift as the key failure mode, while showing that attention allocation plays a limited role in explaining the observed degradation. Overall, we provide a systematic framework for analyzing and improving non-linguistic memory in long-context LALMs, shedding light on future data and training design for robust acoustic memory modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。