arXiv:2502.05489cs.CLcs.AI2025-02ACL被引 31

揭示大模型如何通过特定区域推断情绪,可精准干预生成内容。

Mechanistic Interpretability of Emotion Inference in Large Language Models

  • 发现情绪信息在大模型中集中在特定神经区域。
  • 通过心理评估理论验证其机制合理性,干预后输出符合预期。
  • 适合关注情感生成安全与对齐的研究者使用。

大语言模型(LLMs)在从文本中预测人类情绪方面展现出巨大潜力,但其处理情绪刺激的内在机制仍不明确。本研究通过分析自回归大模型的情绪推理过程,发现情绪表征在模型中具有功能上的局部化特征。实验涵盖多种模型家族与规模,并通过稳健性检验验证结果。进一步结合认知评估理论(cognitive appraisal theory),该理论认为情绪源于对环境刺激的评估,我们通过对构想的评估概念进行因果干预,成功引导模型生成结果,且输出与理论和直觉预期一致。该研究提出了一种新的因果干预方式,可精确控制情感文本生成,对敏感情感领域的安全性与对齐具有潜在应用价值。

原文摘要 · Abstract (English)

Large language models (LLMs) show promising capabilities in predicting human emotions from text. However, the mechanisms through which these models process emotional stimuli remain largely unexplored. Our study addresses this gap by investigating how autoregressive LLMs infer emotions, showing that emotion representations are functionally localized to specific regions in the model. Our evaluation includes diverse model families and sizes and is supported by robustness checks. We then show that the identified representations are psychologically plausible by drawing on cognitive appraisal theory, a well-established psychological framework positing that emotions emerge from evaluations (appraisals) of environmental stimuli. By causally intervening on construed appraisal concepts, we steer the generation and show that the outputs align with theoretical and intuitive expectations. This work highlights a novel way to causally intervene and precisely shape emotional text generation, potentially benefiting safety and alignment in sensitive affective domains.

情绪推理可解释性因果干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。