首次用临床案例验证大模型情绪处理机制,发现情感感知与分类可分离。
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
- 用无关键词的临床情境测试模型情绪识别能力
- 情感感知准确率高达AUROC 1.000,跨模型一致
- 情绪分类依赖关键词,规模越大越不敏感
大语言模型似乎具备情绪内部表征——已有研究报告多个模型家族中存在‘情绪回路’、‘情绪神经元’和结构化情绪流形。但所有研究均使用显式情绪关键词作为刺激,未解之问是:这些回路检测的是真实情感意义,还是仅识别关键词‘devastated’?我们首次基于临床心理学方法,采用去关键词的临床叙事案例(仅通过情境与行为线索引发情绪)对情绪回路假设进行临床有效性检验。在六种模型(Llama-3.2-1B、Llama-3-8B、Gemma-2-9B;base与instruct变体)上应用四种收敛性机制可解释性方法——线性探测、因果激活修补、敲除实验与表示几何分析,发现两种可分离的情绪处理机制:情感接收(检测情感显著内容)在近似完美准确率下运行(AUROC 1.000),表现为早期层饱和,并在全部六种模型中复现;情绪分类(将情感映射至具体标签)部分依赖关键词,去除关键词后性能下降1-7%,且随模型规模提升而改善。因果激活修补证实,含关键词与无关键词刺激共享表示空间,传递的是情感显著性而非情绪类别身份。该结果否定关键词识别假说,建立新型机制分离,引入临床刺激方法作为大模型情绪处理声明的严谨测试标准,对AI安全评估与对齐具有直接意义。所有刺激、代码与数据均已公开供复现。
原文摘要 · Abstract (English)
Large language models appear to develop internal representations of emotion -- "emotion circuits," "emotion neurons," and structured emotional manifolds have been reported across multiple model families. But every study making these claims uses stimuli signalled by explicit emotion keywords, leaving a fundamental question unanswered: do these circuits detect genuine emotional meaning, or do they detect the word "devastated"? We present the first clinical validity test of emotion circuit claims using mechanistic interpretability methods grounded in clinical psychology -- clinical vignettes that evoke emotions through situational and behavioural cues alone, emotion keywords removed. Across six models (Llama-3.2-1B, Llama-3-8B, Gemma-2-9B; base and instruct variants), we apply four convergent mechanistic interpretability methods -- linear probing, causal activation patching, knockout experiments, and representational geometry -- and discover two dissociable emotion processing mechanisms. Affect reception -- detecting emotionally significant content -- operates with near-perfect accuracy (AUROC 1.000), consistent with early-layer saturation, and replicates across all six models. Emotion categorization -- mapping affect to specific emotion labels -- is partially keyword-dependent, dropping 1-7% without keywords and improving with scale. Causal activation patching confirms keyword-rich and keyword-free stimuli share representational space, transferring affective salience rather than emotion-category identity. These findings falsify the keyword-spotting hypothesis, establish a novel mechanistic dissociation, and introduce clinical stimulus methodology as a rigorous standard for testing emotion processing claims in large language models -- with direct implications for AI safety evaluation and alignment. All stimuli, code, and data are released for replication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。