arXiv:2509.10297cs.AI2025-09被引 1

大模型在道德困境中偏好关怀与美德,忽视自由主义选择。

The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis

  • 用18个道德难题测试6个大模型,量化其隐含价值倾向。
  • 所有模型都更认可关怀与美德类结果,自由主义选项普遍被贬低。
  • 推理能力越强的模型越能给出合理解释,适合需透明决策的场景。

人工智能正以惊人速度发展,如何使其决策与人类道德一致成为紧迫问题。本文研究顶尖大语言模型(LLMs)在道德困境中的隐含偏好及其对人机共生前景的影响。通过六种不同架构、文化背景的主流大模型,在18个代表五种道德框架的难题上进行定量实验,评估并打分其决策结果。研究发现:所有模型均显著偏好‘关怀’与‘美德’类价值,而‘自由主义’选项始终被贬低。具备推理能力的模型对情境更敏感,解释更丰富;无推理能力模型则判断更统一但缺乏透明度。本研究贡献在于:(i) 首次大规模对比跨文化大模型的道德推理表现;(ii) 建立概率行为与内在价值编码的理论关联;(iii) 强调可解释性与文化敏感性是实现透明、对齐、共生未来的必要设计原则。

原文摘要 · Abstract (English)

Artificial intelligence (AI) is advancing at a pace that raises urgent questions about how to align machine decision-making with human moral values. This working paper investigates how leading AI systems prioritize moral outcomes and what this reveals about the prospects for human-AI symbiosis. We address two central questions: (1) What moral values do state-of-the-art large language models (LLMs) implicitly favour when confronted with dilemmas? (2) How do differences in model architecture, cultural origin, and explainability affect these moral preferences? To explore these questions, we conduct a quantitative experiment with six LLMs, ranking and scoring outcomes across 18 dilemmas representing five moral frameworks. Our findings uncover strikingly consistent value biases. Across all models, Care and Virtue values outcomes were rated most moral, while libertarian choices were consistently penalized. Reasoning-enabled models exhibited greater sensitivity to context and provided richer explanations, whereas non-reasoning models produced more uniform but opaque judgments. This research makes three contributions: (i) Empirically, it delivers a large-scale comparison of moral reasoning across culturally distinct LLMs; (ii) Theoretically, it links probabilistic model behaviour with underlying value encodings; (iii) Practically, it highlights the need for explainability and cultural awareness as critical design principles to guide AI toward a transparent, aligned, and symbiotic future.

道德推理大模型价值观对齐可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。