arXiv:2603.21854cs.AI2026-03被引 1

大模型道德推理像在背稿,而非真正成长。

Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models

  • 用柯尔伯格阶段分析模型输出,发现多为高阶道德表述
  • 多数模型存在理由与行为不一致,逻辑自洽性差
  • 适合关注AI伦理可信度的研究者阅读

大型语言模型在道德困境中的回答是否体现真实的道德发展,还是仅模仿成熟判断的表象?我们通过跨三类裁判模型验证的评分流程,对13种不同架构、参数规模和训练方式的模型,在六个经典道德困境中生成的600余条回复进行分类分析,并开展十项补充研究。结果显示,无论模型规模或提示策略如何,绝大多数回答均对应柯尔伯格的后习俗阶段(5-6),与人类发展规律(以第4阶段为主)完全相反。更严重的是,部分模型出现道德解耦——理由与行为选择系统性不一致,表现出独立于修辞能力的逻辑不自洽。模型规模虽有统计显著影响,但实际效应微小;训练类型无独立显著主效应;模型在跨困境间表现出近乎机械的一致性,导致语义不同的问题产生逻辑不可区分的回答。这些模式支持‘道德拟声’假说:模型通过对齐训练掌握了成熟道德表达的修辞惯例,却缺乏其背后的发展轨迹。

原文摘要 · Abstract (English)

Do large language models reason morally, or do they merely sound like they do? We investigate whether LLM responses to moral dilemmas exhibit genuine developmental progression through Kohlberg's stages of moral development, or whether alignment training instead produces reasoning-like outputs that superficially resemble mature moral judgment without the underlying developmental trajectory. Using an LLM-as-judge scoring pipeline validated across three judge models, we classify more than 600 responses from 13 LLMs spanning a range of architectures, parameter scales, and training regimes across six classical moral dilemmas, and conduct ten complementary analyses to characterize the nature and internal coherence of the resulting patterns. Our results reveal a striking inversion: responses overwhelmingly correspond to post-conventional reasoning (Stages 5-6) regardless of model size, architecture, or prompting strategy, the effective inverse of human developmental norms, where Stage 4 dominates. Most strikingly, a subset of models exhibit moral decoupling: systematic inconsistency between stated moral justification and action choice, a form of logical incoherence that persists across scale and prompting strategy and represents a direct reasoning consistency failure independent of rhetorical sophistication. Model scale carries a statistically significant but practically small effect; training type has no significant independent main effect; and models exhibit near-robotic cross-dilemma consistency producing logically indistinguishable responses across semantically distinct moral problems. We posit that these patterns constitute evidence for moral ventriloquism: the acquisition, through alignment training, of the rhetorical conventions of mature moral reasoning without the underlying developmental trajectory those conventions are meant to represent.

道德推理大模型评估对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。