arXiv:2608.27402cs.CLcs.AI2026-08

大模型如何组织道德知识?研究发现其在隐空间中以几何方式整合多种道德维度。

How Language Models Organize and Structure Moral Knowledge

论文配图:How Language Models Organize and Structure Moral Knowledge
图 1 · 摘自论文原文
  • 用线性探测器分离六类道德基础,分析其在表示空间中的几何关系
  • 各道德方向共享正向公共成分,体现道德整合性,且早于训练后期出现
  • 模型捕捉道德冲突本身,而非预设判断,适合研究道德认知机制

大型语言模型如何组织道德知识?尽管模型能广泛检测道德内容,但这仅是基础门槛。我们探究它们是否进一步区分不同道德基础,并在表示空间中以几何方式组织其关系。在开源语言模型上训练了六个独立的线性探测器,分别对应道德基础理论(MFT)的六类:关怀/伤害、公平/欺骗、自由/压迫、忠诚/背叛、权威/挑战、神圣/亵渎。研究发现,这些方向既未坍缩为单一检测器,也未完全孤立,而是占据接近最大数量的独立维度,同时共享一个正向公共成分。该公共成分是整合的标志,且相对于同构构建的非道德概念库具有显著特异性(平均成对余弦相似度0.26对比0.013)。该几何结构在不同架构和规模下保持一致,且在预训练早期即达整合状态,远早于探测器准确率饱和。模型所揭示的结构未支持道德基础理论预测的个体化/联结区分(仅20种候选划分),反而反映语料库统计特征。扩展至道德困境,每个困境方向部分由其组成基础构成,相比错配对基线提升2.7倍,但其大部分方差编码的是特定冲突结构。模型表征的是道德张力本身,而非预判的结论。

原文摘要 · Abstract (English)

How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.

道德认知语言模型几何结构表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。