揭示大模型难识别某些情绪的内在原因,提出可干预的因果机制
Why Are Some Emotions Harder for LLMs? Uncovering the Causal Mechanisms of Emotion Inference via Sparse Autoencoders
- 用稀疏自编码器定位情绪推理的因果特征,发现不同情绪特征分布差异
- 惊讶、恐惧特征集中,厌恶特征分散且易被愤怒特征覆盖,导致识别弱
- 可通过定向调整或全局优化特征向量提升模型情绪识别能力
大语言模型在情感敏感的人机交互中广泛应用,但其情绪识别能力参差不齐:对某些情绪表现良好,却持续难以识别其他情绪。尽管已有研究探索模型中的情绪机制,但从可解释性角度理解为何模型在某些情绪上表现更弱仍不清楚。本文通过稀疏自编码器(SAEs)系统分析情绪推理的因果机制,识别出驱动情绪判断的关键稀疏特征,并分析其在情绪内部及跨情绪间的因果组织结构。结果表明,惊讶和恐惧依赖高度集中的特征集,而厌恶则呈现更分散的稀疏因果结构:其因果特征普遍较弱,常与其他情绪特征共激活,且常被愤怒的因果特征所掩盖。这种表征差异为模型在特定情绪上的表现薄弱提供了机制解释。最后,我们进行了两项干预实验:一是针对弱因果特征进行定向调控以缓解情绪特异性失败;二是对识别出的因果特征进行全局优化,显著提升整体情绪识别性能。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in emotionally sensitive human-AI applications, where reliable emotion detection is essential. However, their emotion recognition abilities remain uneven: models often perform well on some emotions while consistently struggling with others. Although recent work has explored emotion mechanisms in LLMs, little is known about why models are weaker on some emotions than others from a mechanistic interpretability perspective. In this work, we investigate emotion-specific biases through the causal mechanisms of emotion inference using sparse autoencoders (SAEs). We systematically identify causal sparse emotion features that drive emotion inference and analyze their sparse causal organization within and across emotions. We show that some emotions, such as surprise and fear, rely on highly concentrated feature sets, whereas disgust exhibits a more distributed sparse causal organization: its causal features are generally weaker, frequently co-activate with features for other emotions, and are often overshadowed by causal features for anger. These representational differences provide a mechanistic explanation for why LLMs struggle more with certain emotions. Finally, we conduct two intervention experiments: targeted steering of weaker causal features to mitigate emotion-specific failures, and global optimization of a steering vector over the identified causal features to improve overall emotion recognition performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。