arXiv:2605.05653cs.CL2026-05

发现大模型处理负面情绪比正面情绪更早,且可定向调控。

Negative Before Positive: Asymmetric Valence Processing in Large Language Models

论文配图:Negative Before Positive: Asymmetric Valence Processing in Large Language Models
图 1 · 摘自论文原文
  • 通过激活修补和控制实验,发现负向情感在早期层处理,正向在中后期层。
  • 固定主题仅改变情感极性时,模型响应反转,排除主题误判可能。
  • 在特定层注入正面信号可让中性输入变正向,证明情感可被直接操控。

机制可解释性揭示了大语言模型(LLMs)如何编码概念,但情感内容的机制理解仍不充分。我们研究了LLMs是否通过专用内部结构或表面标记匹配来处理情感极性。通过对开源LLMs使用激活修补和控制技术,发现负向与正向情感在不同网络深度被处理:负向结果集中在早期层,正向则在中至晚期层达到峰值。在保持主题不变而仅反转情感极性时,模型产生相反响应,排除了主题识别的可能。在识别出的层上使用正面引导方向进行控制,可使中性提示转向正向情感,表明这些层将情感极性编码为可操纵的方向。因此,大模型中的情感极性是局部化、因果性且可控制的,成为基于可解释性的监督目标。

原文摘要 · Abstract (English)

Mechanistic interpretability has revealed how concepts are encoded in large language models (LLMs), but emotional content remains poorly understood at the mechanistic level. We study whether LLMs process emotional valence through dedicated internal structure or through surface token matching. Using activation patching and steering on open-source LLMs, we find that negative and positive valence are processed at different network depths. Negative outcomes localize to early layers while positive outcomes peak at mid-to-late layers. Holding topic fixed while flipping valence produces sign-opposite responses, ruling out topic detection. Steering with the good-news direction at the identified layers shifts neutral prompts toward positive valence, showing these layers encode valence as a manipulable direction. Emotional valence in LLMs is localized, causal and steerable, making it a concrete target for interpretability-based oversight.

情感建模机制解释可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。