arXiv:2508.07284cs.CLcs.AI2025-08被引 4

测试14个大模型在27种道德困境中的决策,发现模型行为受伦理框架影响显著。

"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas

  • 用27种道德情境+10种哲学框架,让模型做二元选择并解释理由
  • 推理增强模型更果断且解释更一致,但未必更符合人类共识
  • 在利他、公平、美德伦理下表现最佳,亲属/法律/自我利益框架易出问题

随着大语言模型(LLMs)越来越多地参与伦理敏感决策,理解其道德推理过程变得至关重要。本研究对14个主流大模型(含推理增强型与通用型)进行了全面实证评估,涵盖27种不同的电车难题场景,基于十种道德哲学框架(包括功利主义、义务论、利他主义等)。通过因子化提示协议,共获取3,780条二元决策及自然语言解释,分析维度包括决策果断性、解释一致性、公众道德对齐度以及对无关伦理线索的敏感性。结果表明,不同伦理框架与模型类型间存在显著差异:推理增强模型表现出更高的决断力和结构化解释,但并不总与人类共识一致。值得注意的是,在利他主义、公平性和美德伦理框架中出现‘甜蜜区’,模型在此类情境下干预率高、解释冲突低、且与聚合人类判断偏差小。然而在强调亲属关系、合法性或自我利益的框架下,模型常产生伦理争议性结果。这些模式表明,道德提示不仅是行为调节工具,更是揭示各厂商模型潜在对齐哲学的诊断手段。研究主张将道德推理作为大模型对齐的核心维度,呼吁建立标准化基准,不仅评估模型‘做什么’,更要考察‘如何做’与‘为何做’。

原文摘要 · Abstract (English)

As large language models (LLMs) increasingly mediate ethically sensitive decisions, understanding their moral reasoning processes becomes imperative. This study presents a comprehensive empirical evaluation of 14 leading LLMs, both reasoning enabled and general purpose, across 27 diverse trolley problem scenarios, framed by ten moral philosophies, including utilitarianism, deontology, and altruism. Using a factorial prompting protocol, we elicited 3,780 binary decisions and natural language justifications, enabling analysis along axes of decisional assertiveness, explanation answer consistency, public moral alignment, and sensitivity to ethically irrelevant cues. Our findings reveal significant variability across ethical frames and model types: reasoning enhanced models demonstrate greater decisiveness and structured justifications, yet do not always align better with human consensus. Notably, "sweet zones" emerge in altruistic, fairness, and virtue ethics framings, where models achieve a balance of high intervention rates, low explanation conflict, and minimal divergence from aggregated human judgments. However, models diverge under frames emphasizing kinship, legality, or self interest, often producing ethically controversial outcomes. These patterns suggest that moral prompting is not only a behavioral modifier but also a diagnostic tool for uncovering latent alignment philosophies across providers. We advocate for moral reasoning to become a primary axis in LLM alignment, calling for standardized benchmarks that evaluate not just what LLMs decide, but how and why.

道德推理大模型对齐伦理框架电车难题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。