arXiv:2608.17102cs.CLeess.AS2026-08

发现语音与表情情绪识别共享稀疏解码器组件,可无训练操控。

Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

论文配图:Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
图 1 · 摘自论文原文
  • 通过情绪敏感神经元定位跨模态情感表征
  • 视觉神经元失活导致对应表情识别下降
  • 跨模态干预实现双向情绪调控,适合多模态研究者

现代多模态基础模型(MFMs)在整合语音、视觉和语言感知的任务中进展迅速,包括情绪识别。然而,它们是通过共享的情感功能单元还是特定模态路径来识别语音与面部情绪仍不明确。本文在三个MFMs——Gemma-4-12B-it、MiniCPM-o-4.5和Qwen2.5-Omni-7B中探索情绪敏感神经元(ESNs),即与情绪类别显著关联的稀疏解码器神经元。利用语音情绪识别与面部表情识别作为互补探针,识别出声学与视觉的ESNs。视觉ESNs具有因果意义:关闭它们会特异性损害对应面部情绪的识别,而调节其激活可选择性增强该情绪相对于其他情绪的识别。声学与视觉ESNs还表现出情绪匹配的重叠及相似的层间分布,表明跨模态情感表征存在部分结构对齐。最后,跨模态干预揭示双向因果转移:来自一模态的ESNs在另一模态上施加情绪特异性影响。本研究为MFMs中情感功能单元的跨模态激活级分析提供了首批证据,表明语音与面部情绪识别部分汇聚于可定位与操控的稀疏解码器层级组件,无需再训练。

原文摘要 · Abstract (English)

Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways. We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs. Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories. Acoustic and visual ESNs further show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment between affective representations across speech and faces. Finally, cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other. Our findings provide one of the first cross-modality activation-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder-level components that can be localized and manipulated without training.

多模态情绪识别神经元可解释性稀疏解码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。