用五阶段道德困境测试大模型的动态伦理判断能力
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
- 构建3302个渐进式道德困境,评估模型随情境变化的伦理推理
- 九款大模型在复杂情境中价值偏好明显转变,体现动态调整
- 发现关怀优先但公平可超越关怀,适合关注伦理对齐的研究者
伦理决策是人类判断的关键方面,随着大语言模型在决策支持系统中的广泛应用,对其道德推理能力的严格评估变得尤为重要。然而,现有评估多依赖单步测试,难以捕捉模型应对不断演变的伦理挑战时的表现。为此,我们提出多步道德困境(MMDs)数据集,首次专门设计用于评估大语言模型在3302个五阶段困境中逐步演变的道德判断。该框架实现了对模型道德推理动态演化的细粒度分析。我们对九款广泛使用的大型语言模型的评估显示,其价值偏好在困境推进过程中发生显著变化,表明模型会根据情景复杂性重新校准道德判断。此外,成对价值比较表明,尽管模型通常优先考虑关怀价值,但在特定情境下公平价值可能超越关怀,凸显了大模型伦理推理的动态与情境依赖性。研究呼吁转向动态、情境感知的评估范式,为更符合人类价值观、更具敏感性的大模型发展铺平道路。
原文摘要 · Abstract (English)
Ethical decision-making is a critical aspect of human judgment, and the growing use of LLMs in decision-support systems necessitates a rigorous evaluation of their moral reasoning capabilities. However, existing assessments primarily rely on single-step evaluations, failing to capture how models adapt to evolving ethical challenges. Addressing this gap, we introduce the Multi-step Moral Dilemmas (MMDs), the first dataset specifically constructed to evaluate the evolving moral judgments of LLMs across 3,302 five-stage dilemmas. This framework enables a fine-grained, dynamic analysis of how LLMs adjust their moral reasoning across escalating dilemmas. Our evaluation of nine widely used LLMs reveals that their value preferences shift significantly as dilemmas progress, indicating that models recalibrate moral judgments based on scenario complexity. Furthermore, pairwise value comparisons demonstrate that while LLMs often prioritize the value of care, this value can sometimes be superseded by fairness in certain contexts, highlighting the dynamic and context-dependent nature of LLM ethical reasoning. Our findings call for a shift toward dynamic, context-aware evaluation paradigms, paving the way for more human-aligned and value-sensitive development of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。