arXiv:2607.11801cs.SDcs.AI2026-07

通过精准放大音频编码器中的关键神经元,显著提升大模型对语音情感等细节的感知能力。

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

论文配图:Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
图 1 · 摘自论文原文
  • 在音频编码器内识别并放大对声学特征敏感的神经元,无需训练
  • 在10项非语义语音属性上平均准确率提升超20点,最高达25.7
  • 仅干预编码器内部神经元即有效,解码侧或语言模型内干预效果差

大型音频-语言模型(LALMs)在语音内容理解上表现良好,但在说话人情绪等细粒度非语义属性上表现不足。现有方法多在音频编码后干预,且粒度较粗。本文提出IAAN——一种无需训练、无标签的推理时干预方法,通过对比真实波形与噪声参考下各前馈神经元的激活值,评估其对声学信息的敏感度,并在推理时放大得分最高的少量神经元。在十项非语义语音属性任务中,IAAN使Audio-Flamingo-3平均准确率提升25.7点,Qwen2.5-Omni提升21.4点,Kimi-Audio提升9.7点。即使在已专门微调以重视声学证据的模型上仍有效。控制实验表明,编码器位置和神经元级选择性均不可或缺;在解码侧或语言模型内干预几乎无效,甚至降低性能。提升效果依赖于具体被放大的神经元,而非数量,证明了声学评分的有效性。结果表明,对音频编码器内部神经元进行小规模精准干预,是提升大模型声学理解力的新路径。

原文摘要 · Abstract (English)

Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content. Improving this without the cost of retraining calls for an effective inference-time intervention, yet most existing methods intervene only after the audio encoder and operate at a relatively coarse granularity. The encoder itself, where acoustic information is first extracted from the waveform, remains largely unexplored, especially at the level of individual neurons. We introduce IAAN, Identifying and Amplifying Acoustic Neurons, a training-free and label-free method that scores each feed-forward neuron in the audio encoder by contrasting its activation on the real waveform with that on a noise reference lacking the real audio's acoustic information. IAAN then amplifies a small set of the highest-scoring neurons at inference. Across ten non-semantic speech attributes, IAAN improves average accuracy by 25.7 points on Audio-Flamingo-3, 21.4 on Qwen2.5-Omni, and 9.7 on Kimi-Audio. It also improves a model already explicitly fine-tuned to prioritize acoustic evidence. In controlled comparisons, both the encoder locus and neuron-level selectivity prove necessary for this gain. Intervening after the encoder, at the decoding side or inside the language model, yields little to no improvement, or even deteriorates accuracy. The improvement also depends on which specific neurons are amplified, not merely on their number, confirming that IAAN's acoustic score succeeds in identifying the neurons that matter. These results show that a small, precisely targeted intervention inside the audio encoder is an effective and largely untapped way to strengthen the acoustic understanding of LALMs, opening a new direction for inference-time methods that improve acoustic perception through neuron-level access to the encoder.

音频理解神经元干预大模型优化声学感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。