arXiv:2601.05437cs.CLcs.AI2026-01被引 4

揭示大模型内在道德结构如何从语言统计中自然浮现

Tracing Moral Foundations in Large Language Models

  • 用道德基础理论分析14个模型的道德表征层次
  • 发现模型对道德概念的区分与人类判断高度一致
  • 支持道德输出的神经机制可被精准干预和引导

大语言模型常生成类人道德判断,但其背后是深层概念结构还是表面模仿尚不明确。本文以道德基础理论(MFT)为框架,研究了14个基础及指令微调模型(涵盖Llama、Qwen2.5、Qwen3-MoE、Mistral四大家族,参数量7B至70B)中道德基础的编码、组织与表达方式。采用多层次方法:(i)逐层分析MFT概念表征与人类道德感知的一致性;(ii)利用预训练稀疏自编码器(SAEs)在残差流中识别支持道德概念的稀疏特征;(iii)通过密集MFT向量与稀疏SAE特征进行因果干预。结果表明,模型对道德基础的表征与人类判断高度一致,且这种道德几何结构由预训练自然形成,后训练阶段可选择性重构。细粒度分析显示,SAE特征与特定道德基础存在清晰语义关联,表明共享表示中存在部分解耦机制。无论使用密集向量或稀疏特征进行干预,均能预测性地改变与道德基础相关的行为,证明内部表征与道德输出间存在因果联系。整体表明,大模型中的多元道德结构可作为语言统计规律的潜在模式自然涌现。

原文摘要 · Abstract (English)

Large language models often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or superficial ``moral mimicry.'' Using Moral Foundations Theory (MFT) as an analytic framework, we study how moral foundations are encoded, organized, and expressed across 14 base and instruction-tuned LLMs spanning four model families (Llama, Qwen2.5, Qwen3-MoE, Mistral) and scales from 7B to 70B. We employ a multi-level approach combining (i) layer-wise analysis of MFT concept representations and their alignment with human moral perceptions, (ii) pretrained sparse autoencoders (SAEs) over the residual stream to identify sparse features that support moral concepts, and (iii) causal steering interventions using dense MFT vectors and sparse SAE features. We find that models represent and distinguish moral foundations in a manner that aligns with human judgments, and that this moral geometry naturally emerges from pretraining and is selectively rewired by post-training. At a finer scale, SAE features show clear semantic links to specific foundations, suggesting partially disentangled mechanisms within shared representations. Finally, steering along either dense vectors or sparse features produces predictable shifts in foundation-relevant behavior, demonstrating a causal connection between internal representations and moral outputs. Together, our results provide mechanistic evidence that moral concepts in LLMs are distributed, layered, and partly disentangled, suggesting that pluralistic moral structure can emerge as a latent pattern from the statistical regularities of language alone.

大模型伦理道德认知表征分析因果干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。