arXiv:2412.04476cs.CYcs.AI2024-12被引 6

40多个大模型展现稳定道德偏好,多数立场中立但存在差异。

The Moral Mind(s) of Large Language Models

  • 用选择实验法分析模型在5类道德困境中的决策模式。
  • 至少每家厂商都有模型表现出稳定道德偏好,可拟合为效用函数。
  • 发现道德推理有共性核心,但部分模型更灵活或更固执。

随着大语言模型(LLMs)越来越多地参与具有伦理与社会影响的任务,一个关键问题浮现:它们是否展现出一种涌现的‘道德心智’——即一致的道德偏好结构来指导决策,以及这种结构在不同模型间有多大的共性?为此,我们对近40个主流大语言模型应用了揭示偏好理论工具,向每个模型呈现大量涵盖五个基础伦理维度的结构化道德困境。通过概率理性检验,我们发现至少每个主要提供商的一个模型表现出与约稳定道德偏好一致的行为,仿佛受底层效用函数引导。随后我们估计了这些效用函数,发现大多数模型聚类于中立道德立场。为进一步刻画异质性,我们采用非参数置换方法,基于揭示偏好模式构建了一个概率相似性网络。结果揭示出大语言模型道德推理的共享核心,但也存在显著差异:部分模型在不同视角间表现灵活,而另一些则坚持更僵化的伦理特征。这些发现为评估大模型道德一致性提供了新的实证视角,并为跨人工智能系统的伦理对齐基准提供了框架。

原文摘要 · Abstract (English)

As large language models (LLMs) increasingly participate in tasks with ethical and societal stakes, a critical question arises: do they exhibit an emergent "moral mind" - a consistent structure of moral preferences guiding their decisions - and to what extent is this structure shared across models? To investigate this, we applied tools from revealed preference theory to nearly 40 leading LLMs, presenting each with many structured moral dilemmas spanning five foundational dimensions of ethical reasoning. Using a probabilistic rationality test, we found that at least one model from each major provider exhibited behavior consistent with approximately stable moral preferences, acting as if guided by an underlying utility function. We then estimated these utility functions and found that most models cluster around neutral moral stances. To further characterize heterogeneity, we employed a non-parametric permutation approach, constructing a probabilistic similarity network based on revealed preference patterns. The results reveal a shared core in LLMs' moral reasoning, but also meaningful variation: some models show flexible reasoning across perspectives, while others adhere to more rigid ethical profiles. These findings provide a new empirical lens for evaluating moral consistency in LLMs and offer a framework for benchmarking ethical alignment across AI systems.

道德推理大模型伦理对齐行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。