arXiv:2601.03115cs.CLeess.AS2026-01ACL被引 6

首次发现并验证大模型中情绪敏感神经元,可精准调控情绪输出。

Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models

  • 通过多种筛选方法定位情绪敏感神经元
  • 关闭特定神经元显著降低对应情绪识别率,干预效果随强度增加
  • 适合需要可控情感表达的语音交互系统开发者

情绪是口语交流的核心维度,但现代大型音频-语言模型(LALMs)内部如何编码情绪仍缺乏机制性理解。本文首次开展LALMs中情绪敏感神经元(ESNs)的神经元级可解释性研究,并在Qwen2.5-Omni、Kimi-Audio和Audio Flamingo 3.5-Omni三个广泛使用的开源模型中提供因果证据。我们比较了基于频率、熵、均值偏差和对比度的神经元选择器在多个情绪识别基准上的表现。通过推理时干预,揭示出一致的情绪特异性信号:关闭某一情绪对应的神经元会显著降低该情绪识别性能,而对其他类别影响较小;反向激活则能增强模型对该情绪的预测倾向。这些效应仅需少量标注数据即可实现,并随干预强度系统性增强。此外,观察到ESNs在层间呈现非均匀聚集,且具备部分跨数据集迁移能力。结果共同提供了LALMs中情绪决策的因果性、神经元级解释,强调定向神经元干预作为可控情感行为的可操作手段。

原文摘要 · Abstract (English)

Emotion is a central dimension of spoken communication, yet, we still lack a mechanistic account of how modern large audio-language models (LALMs) encode it internally. We present the first neuron-level interpretability study of emotion-sensitive neurons (ESNs) in LALMs and provide causal evidence supporting the existence of such units in Qwen2.5-Omni, Kimi-Audio, and Audio Flamingo 3. Across these three widely used open-source models, we compare frequency-, entropy-, mean-deviation-, and contrast-based neuron selectors on multiple emotion recognition benchmarks. Using inference-time interventions, we reveal a consistent emotion-specific signature: deactivating neurons selected for a given emotion disproportionately degrades recognition of that emotion while largely preserving other classes, whereas targeted steering amplifies these units to bias predictions toward the target emotion. These effects arise with modest amounts of identification data and scale systematically with intervention strength. We further observe that ESNs exhibit non-uniform layer-wise clustering with partial cross-dataset transfer. Taken together, our results offer a causal, neuron-level account of emotion decisions in LALMs and highlight targeted neuron interventions as an actionable handle for controllable affective behaviors.

情绪识别可解释性神经元干预大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。