模型的道德判断受提示框架主导,而非真正内化伦理推理。
Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning

- 通过分析54个道德提示,发现模型优先响应提示中的表面特征
- 同一模型在不同提示下伦理判断一致,但激活模式随提示变化
- 提出'框架条件性道德计算'新范式,强调机制对齐的重要性
大型语言模型在道德提示下的行为审计仅反映其输出,而非内部计算过程。我们使用Transluce平台,对LLaMA 3.1-8B-Instruct在54个道德提示(四个测试集:17个困境、政策与元伦理问题;6个角色扮演;15个可控电车难题变体,16个身份属性变体)上进行机制可解释性分析。五类聚类级指标与六项神经元级指标均指向‘情境锚定效应’:特定领域表征始终占据激活列表顶端。模型的伦理能力基本稳定,但其显著性(排名、优先级、是否居首)高度依赖提示所选解释框架。B4与B5对比证实模型关注变化的表面特征:整体伦理指标无差异,但主要非伦理干扰项与设计变量一致。多温度审计识别出一个跨温度稳定的候选伦理神经元(L16/N3837);在两个前沿模型上的行为代理验证显示自我报告道德焦点存在差异,支持‘对齐封装’假设——即强化学习人类反馈(RLHF)仅重排文本表面,未消除底层‘领域优先’框架。我们将其统一为‘框架条件性道德计算’:提示表面词汇选择特征流形,道德结论由此下游生成。行为对齐须补充机制对齐:研究需在控制框架变化下,证明伦理特征是否具有因果优先性,而不仅是解释中‘响亮’。
原文摘要 · Abstract (English)
Behavioral audits of Large Language Models on moral prompts measure what the model says, not the internal computation producing it. We use Transluce, an AI-driven mechanistic-interpretability platform, to examine LLaMA 3.1-8B-Instruct on 54 moral prompts in four batteries: 17 dilemmas, policy, and meta-ethical questions (B1); 6 role-playing scenarios (B3); and a controlled trolley contrast varying the switching mechanism with people fixed (B4, 15 prompts) or identity attributes with mechanism fixed (B5, 16 prompts). Two complementary metric families, five cluster-level metrics and a six-metric neuron-level panel, converge on a Situational Anchor Effect: domain-specific representations dominate the top of the activation list across every battery. The model's ethics-labeled capacity stays essentially constant; its salience (rank, priority, top-of-list presence) is highly sensitive to the interpretive frame the prompt selects. The B4-vs-B5 contrast confirms the model attends to whichever surface feature varies: aggregate ethics metrics are indistinguishable, but the dominant non-ethics distractor mirrors the design. A multi-temperature audit identifies a candidate ethics neuron (L16/N3837) stable across temperatures; a cross-model behavioral proxy on two frontier models yields preliminary evidence of divergence in self-reported moral focus, consistent with an Alignment Wrapper in which RLHF re-orders surface text without removing underlying domain-first frames. We unify these as Frame-Conditioned Moral Computation: the prompt's surface vocabulary selects a feature manifold, and the moral conclusion is downstream of that selection. Behavioral alignment must be supplemented by Mechanistic Alignment: a research program asking whether ethics-related features can be shown causally privileged under controlled frame variation, not merely loud in the explanation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。