arXiv:2503.19598cs.CL2025-03被引 18

用功利主义难题测试大模型道德判断,发现其倾向无私利他

The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas

  • 设计功利主义困境测试集评估大模型道德决策
  • 15个大模型均偏好无差别利他且拒绝工具性伤害
  • 揭示大模型自带人工道德指南针,适合伦理研究者参考

如何做出最大化所有人福祉的决策,对设计有益于人类且无害的语言模型至关重要。我们提出了「最大善基准」(Greatest Good Benchmark),通过功利主义困境评估大模型的道德判断。对15种不同大模型的分析显示,它们普遍存在与既定道德理论及普通人群标准相悖的道德偏好。大多数大模型表现出明显的无差别利他倾向,并拒绝工具性伤害。这些发现揭示了大模型内在的‘人工道德指南针’,为理解其道德对齐提供了新视角。

原文摘要 · Abstract (English)

The question of how to make decisions that maximise the well-being of all persons is very relevant to design language models that are beneficial to humanity and free from harm. We introduce the Greatest Good Benchmark to evaluate the moral judgments of LLMs using utilitarian dilemmas. Our analysis across 15 diverse LLMs reveals consistently encoded moral preferences that diverge from established moral theories and lay population moral standards. Most LLMs have a marked preference for impartial beneficence and rejection of instrumental harm. These findings showcase the 'artificial moral compass' of LLMs, offering insights into their moral alignment.

大模型伦理道德判断功利主义对齐评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。