用无关键词故事测试语言模型情绪机制,避免词语干扰。
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models

- 设计480个无情绪词的故事,通过情节引发八种基本情绪
- 每条情绪故事配一个情绪消除的对照故事,确保对比公平
- 适合研究模型情绪表征、可解释性或对抗性测试的学者
大语言模型情绪机制研究(如线性探测、激活修补、稀疏自编码器分析、因果消融、引导向量提取)依赖包含情绪词汇的刺激材料。当模型对“我非常愤怒”产生响应时,无法判断是识别了愤怒情绪,还是仅识别了“愤怒”一词。这会严重影响关于情绪回路、特征与干预措施的结论可信度。为此,我们发布AIPsy-Affect,一个480项临床刺激库:192个无关键词叙事片段,分别诱发普拉切克八种基本情绪;192个匹配的中性对照,人物、背景、长度和表面结构一致,但情绪被手术式移除;另设中等强度与判别效度划分。成对结构提供强方法学保障:任何能区分临床项与中性对照的内部表征,不可能基于情绪关键词存在。三种自然语言处理防御测试(词袋情感分析、情绪词典、上下文变换器分类器)验证该性质:词袋方法仅感知情境词汇,而上下文分类器虽能检测情绪(p < 10^-15),却无法识别类别(准确率5.2% vs. 关键词对照组82.5%)。AIPsy-Affect在先前96项电池基础上扩展四倍,开源发布,许可证为MIT。
原文摘要 · Abstract (English)
Mechanistic interpretability research on emotion in large language models -- linear probing, activation patching, sparse autoencoder (SAE) feature analysis, causal ablation, steering vector extraction -- depends on stimuli that contain the words for the emotions they test. When a probe fires on "I am furious", it is unclear whether the model has detected anger or detected the word "furious". The two readings have very different consequences for every downstream claim about emotion circuits, features, and interventions. We release AIPsy-Affect, a 480-item clinical stimulus battery that removes the confound at the stimulus level: 192 keyword-free vignettes evoking each of Plutchik's eight primary emotions through narrative situation alone, 192 matched neutral controls that share characters, setting, length, and surface structure with the affect surgically removed, plus moderate-intensity and discriminant-validity splits. The matched-pair structure supports linear probing, activation patching, SAE feature analysis, causal ablation, and steering vector extraction under a strong methodological guarantee: any internal representation that distinguishes a clinical item from its matched neutral cannot be doing so on the basis of emotion-keyword presence. A three-method NLP defense battery -- bag-of-words sentiment, an emotion-category lexicon, and a contextual transformer classifier -- confirms the property: bag-of-words methods see only situational vocabulary, and a contextual classifier detects affect (p < 10^-15) but cannot identify the category (5.2% top-1 vs. 82.5% on a keyword-rich control). AIPsy-Affect extends our earlier 96-item battery (arXiv:2603.22295) by a factor of four and is released openly under MIT license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。