arXiv:2608.08061cs.AI2026-08中稿 · the Fourth Interna…

测试大模型在道德冲突中按优先级权衡伤害的能力,发现多数模型默认回避直接伤害。

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

论文配图:CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 设计90个层级化道德困境,测试模型在不同伦理框架下的优先级判断
  • 9/10模型优先避免直接伤害,而非最小化总体损失
  • 揭示模型难遵循指定优先级,如人>动物>机器人,适合评估伦理可控性

当前大模型的道德评估多聚焦于答案是否符合道德或避免明显错误,而非在无完美选项时如何权衡冲突原则。本文提出CORDA(条件排序与分级指令遵从)基准,基于道德链形式化,评估大模型在90个道德困境中的层级化、以伤害为中心的推理能力,涵盖电车难题、医疗权衡、资源分配及人-动物-机器人冲突,涉及四个有序伦理框架:效用、效用+主体伤害、双过程、双过程+主体伤害。测试显示,10个指令微调模型中有9个表现出强烈的义务论倾向,优先避免直接人身伤害;模型在识别道德红线(如不杀人)方面表现更可靠,而在基于结果的比较(如最小化总伤害)上较弱。尽管所有模型能响应显式条件,但多个模型无法一致遵循指定优先级(如人类>动物>机器人)。CORDA填补了大模型道德评估的关键空白,强调道德可靠性不仅需默认克制,更需在冲突中可控地应用上下文指定的优先级。

原文摘要 · Abstract (English)

The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluations of large language models (LLMs) remain limited: most test whether models give morally acceptable answers, match human preferences, or avoid obvious violations, rather than whether they can prioritise between competing principles when no option is morally cost-free. We introduce CORDA (Conditioned Ordering and Ranked Directive Adherence), a benchmark for evaluating hierarchical, harm-centred moral reasoning in LLMs. Building on the morality chains formalism, CORDA tests 90 moral dilemmas involving trolley-style cases, medical trade-offs, resource allocation, and human-animal-robot conflicts across four ordered ethical frameworks: Utility, Utility + Agent Harm, Dual-Process, and Dual-Process + Agent Harm. Together, these frameworks test whether models can adapt their decisions when moral priorities change. Across ten instruction-tuned models from seven providers, we find a strong deontological default, with 9 of 10 prioritising avoidance of direct personal harm over reducing overall harm. Models also perform more reliably on categorical harm-avoidance rules, such as avoiding killing, than on outcome-based comparisons, such as minimising total harm, suggesting that they recognise moral red lines more easily than they reason through competing harms. Although all models respond to explicit chain conditioning, several fail to consistently follow specified priority orderings, such as humans over animals and animals over robots. CORDA addresses a central gap in LLM moral evaluation by testing whether models can move beyond default harm-avoidant responses and apply context-specified moral priorities. Moral reliability requires more than default restraint; it requires controllability under conflict.

道德推理大模型评估伦理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。