揭示大模型在痛乐决策中的内部机制,找到关键计算位置。
Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM
- 通过分层探查与激活干预,定位痛乐信息的内部表示位置。
- 晚期注意力输出处的值对决策影响最大,且强度呈剂量效应。
- 多头协同作用,非单一模块主导,适合安全与伦理研究者参考。
先前行为研究表明,部分大模型在选项被表述为引起痛苦或愉悦时会改变选择,且这种偏差随强度表述增大而增强。为连接行为表现(模型做了什么)与可解释性机制(内部如何计算),本文以Gemma-2-9B-it为基础,采用简化决策任务,系统分析:(i) 通过分层线性探查,映射不同路径中表征的可用性;(ii) 利用激活干预(引导、插补/消融)测试因果贡献;(iii) 在epsilon网格上量化剂量反应,读出2-3类别的逻辑差值及归一化选择概率。结果发现:(a) 痛乐极性在早期层(L0-L1)即可完美线性分离,词法基线仍保留显著信号;(b) 强度分级信息在中后期层强可解码,尤其集中在注意力与MLP输出,决策对齐最高出现在最终标记前;(c) 沿数据驱动的痛乐方向进行叠加引导,在晚期节点显著调节2-3逻辑差值,效果最显著于晚层注意力输出(attn_out L14);(d) 头级别插补/消融表明,此类效应分布于多个头而非集中于单一单元。这些结果将行为敏感性与可识别的内部表征和可干预位置联系起来,为更严格的反事实检验和广泛复现提供具体机制靶点。本工作支持以实证为基础,推动关于AI意识与福祉的讨论,以及政策制定、审计标准与安全防护的治理框架。
原文摘要 · Abstract (English)
Prior behavioural work suggests that some LLMs alter choices when options are framed as causing pain or pleasure, and that such deviations can scale with stated intensity. To bridge behavioural evidence (what the model does) with mechanistic interpretability (what computations support it), we investigate how valence-related information is represented and where it is causally used inside a transformer. Using Gemma-2-9B-it and a minimalist decision task modelled on prior work, we (i) map representational availability with layer-wise linear probing across streams, (ii) test causal contribution with activation interventions (steering; patching/ablation), and (iii) quantify dose-response effects over an epsilon grid, reading out both the 2-3 logit margin and digit-pair-normalised choice probabilities. We find that (a) valence sign (pain vs. pleasure) is perfectly linearly separable across stream families from very early layers (L0-L1), while a lexical baseline retains substantial signal; (b) graded intensity is strongly decodable, with peaks in mid-to-late layers and especially in attention/MLP outputs, and decision alignment is highest slightly before the final token; (c) additive steering along a data-derived valence direction causally modulates the 2-3 margin at late sites, with the largest effects observed in late-layer attention outputs (attn_out L14); and (d) head-level patching/ablation suggests that these effects are distributed across multiple heads rather than concentrated in a single unit. Together, these results link behavioural sensitivity to identifiable internal representations and intervention-sensitive sites, providing concrete mechanistic targets for more stringent counterfactual tests and broader replication. This work supports a more evidence-driven (a) debate on AI sentience and welfare, and (b) governance when setting policy, auditing standards, and safety safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。