无需训练的听觉注意力架构,让模型只关注关键声音事件
NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

- 用类脑振荡工作记忆识别声音中的显著性信号
- 在XD-Violence上将准确率从53.5%提升至70.6%
- 适合处理长音频中稀有事件检测任务
音频提供关键情境线索,但现有音频语言模型在长录音中因背景噪声主导而难以捕捉罕见显著事件。本文提出NAACA,一种无需训练的神经听觉注意认知架构,将注意力分配重构为听觉显著性过滤问题。核心是受脑启发的振荡工作记忆(OWM),能维持稳定吸引子状态,在自适应能量波动提示感知显著性时才触发高层语言模型处理,避免无效调用。在XD-Violence数据集上,NAACA将AudioQwen的平均精度(AP)从53.50%提升至70.60%,显著减少不必要的模型调用。在世界城市声景(USoW)数据集的定性案例中,OWM成功捕捉新事件与子类别变化,且对短暂停顿和环境噪声保持鲁棒。
原文摘要 · Abstract (English)
Audio provides critical situational cues, yet current Audio Language Models (ALMs) face an attention bottleneck in long-form recordings where dominant background patterns can dilute rare, salient events. We introduce NAACA, a training-free NeuroAuditory Attentive Cognitive Architecture that reframes attention allocation as an auditory salience filtering problem. At its core is OWM, a neuro-inspired Oscillatory Working Memory that maintains stable attractor-like states and triggers higher-cognition ALM processing only when adaptive energy fluctuations signal perceptual salience, triggering higher-level reasoning. On XD-Violence, NAACA improves AudioQwen's average precision (AP) from 53.50% to 70.60% while reducing unnecessary ALM invocations. Furthermore, qualitative case studies on the Urban Soundscapes of the World (USoW) dataset show that OWM captures novel events and subcategory shifts while remaining robust to transient pauses and ambient urban noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。