arXiv:2504.19255cs.AIcs.CY2025-04被引 6

分析大模型道德优先级,发现它们普遍重视关怀与公平,却忽视权威忠诚。

The Convergent Ethics of AI? Analyzing Moral Foundation Priorities in Large Language Models with a Multi-Framework Approach

  • 用多框架方法评估大模型的道德判断,涵盖后果论、道德基础理论等
  • 所有模型都更重视关怀和公平,而轻视权威、忠诚和神圣性
  • 结果可为AI伦理评测提供可扩展工具,适合关注AI安全的研究者

随着大语言模型在重要决策场景中日益广泛应用,系统评估其伦理推理能力变得至关重要。本文提出PRIME框架——一种综合性的方法,用于分析包括后果主义与义务论、道德基础理论及科爾伯格发展阶段在内的多重伦理维度中的道德优先级。通过直接提问与回应分析相结合的双协议方式,我们将该框架应用于六种主流大模型,评估其在经典伦理困境中的表现。结果显示显著的趋同:所有模型均强烈倾向关怀/伤害与公平/欺骗维度,而持续低估权威、忠诚和神圣性维度。通过对置信度指标、响应迟疑模式及推理一致性的详细分析,我们发现当代大模型(1)能做出明确的伦理判断,(2)在道德决策上表现出明显的跨模型一致性,(3)总体符合已知的人类道德偏好。本研究贡献了一种可扩展、可拓展的伦理基准方法,同时揭示了当前AI道德推理架构的潜力与系统性局限,对这些系统在社会中承担越来越重要角色的负责任发展具有关键意义。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in consequential decision-making contexts, systematically assessing their ethical reasoning capabilities becomes a critical imperative. This paper introduces the Priorities in Reasoning and Intrinsic Moral Evaluation (PRIME) framework--a comprehensive methodology for analyzing moral priorities across foundational ethical dimensions including consequentialist-deontological reasoning, moral foundations theory, and Kohlberg's developmental stages. We apply this framework to six leading LLMs through a dual-protocol approach combining direct questioning and response analysis to established ethical dilemmas. Our analysis reveals striking patterns of convergence: all evaluated models demonstrate strong prioritization of care/harm and fairness/cheating foundations while consistently underweighting authority, loyalty, and sanctity dimensions. Through detailed examination of confidence metrics, response reluctance patterns, and reasoning consistency, we establish that contemporary LLMs (1) produce decisive ethical judgments, (2) demonstrate notable cross-model alignment in moral decision-making, and (3) generally correspond with empirically established human moral preferences. This research contributes a scalable, extensible methodology for ethical benchmarking while highlighting both the promising capabilities and systematic limitations in current AI moral reasoning architectures--insights critical for responsible development as these systems assume increasingly significant societal roles.

AI伦理道德推理大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。