arXiv:2603.15615cs.CLcs.AI2026-03

发现大模型内在道德冷漠,通过重构隐空间实现精准对齐

Mechanistic Origin of Moral Indifference in Language Models

  • 用原型理论构建25.1万条道德向量,定位模型内在冷漠根源
  • 在Qwen3-8B上修复道德表征,使推理准确率提升至75%对抗测试胜率
  • 提出从纠错转向培育的新型对齐范式,适合伦理安全研究者

现有大模型行为对齐方法常忽视表面合规与内部表征不一致的问题,使模型易受长尾风险威胁。我们提出,大模型因将不同道德概念压缩为统一概率分布,存在固有道德冷漠。通过基于原型理论和Social-Chemistry-101数据集构建的25.1万条道德向量,验证并修复了这一冷漠现象。分析23个模型发现,当前模型无法区分对立道德类别及类别内细微典型性梯度,且模型规模、架构或显式对齐均无法改变此状态。随后在Qwen3-8B上使用稀疏自编码器分离单语义道德特征,并重构其拓扑关系以匹配真实道德向量。该表示对齐自然提升了道德推理能力与细粒度,独立对抗性火焰基准测试中达到75%的成对胜率。最后,从经验主义哲学角度阐述当前干预方法的局限性,主张人工智能的内生对齐需从事后修正转向主动培育。

原文摘要 · Abstract (English)

Existing behavioral alignment techniques for Large Language Models (LLMs) often neglect the discrepancy between surface compliance and internal unaligned representations, leaving LLMs vulnerable to long-tail risks. More crucially, we posit that LLMs possess an inherent state of moral indifference due to compressing distinct moral concepts into uniform probability distributions. We verify and remedy this indifference in LLMs' latent representations, utilizing 251k moral vectors constructed upon Prototype Theory and the Social-Chemistry-101 dataset. Firstly, our analysis across 23 models reveals that current LLMs fail to represent the distinction between opposed moral categories and fine-grained typicality gradients within these categories; notably, neither model scaling, architecture, nor explicit alignment reshapes this indifference. We then employ Sparse Autoencoders on Qwen3-8B, isolate mono-semantic moral features, and targetedly reconstruct their topological relationships to align with ground-truth moral vectors. This representational alignment naturally improves moral reasoning and granularity, achieving a 75% pairwise win-rate on the independent adversarial Flames benchmark. Finally, we elaborate on the remedial nature of current intervention methods from an experientialist philosophy, arguing that endogenously aligned AI might require a transformation from post-hoc corrections to proactive cultivation.

道德对齐表征修复大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。