用真实生活道德困境测试大模型,发现其判断与人类差异大。
Normative Evaluation of Large Language Models with Everyday Moral Dilemmas
- 用Reddit AITA社区1万+真实道德难题评估模型
- 模型判断与人类差异显著,自洽但彼此不一致
- 适合关注AI伦理、决策偏差的研究者
大语言模型(LLMs)的快速应用推动了对其内在道德规范与决策过程的研究。现有研究多通过问卷式提问评估模型与特定群体、道德观念或政治立场的对齐程度,但此类方法往往简化了日常道德困境的复杂性与细微差别。我们主张,应基于更细致的人类互动维度来审计LLMs,以更准确评估其对人类信念与行为的影响。为此,我们选取Reddit的“Am I the Asshole”(AITA)社区中超过10,000个真实生活道德冲突案例,让七种主流LLMs判断责任归属并提供解释。随后将模型判断与红迪网友及模型间进行对比,揭示其道德推理模式。结果表明,大模型表现出显著不同的道德判断模式,与人类在AITA社区的评价存在明显差异;模型具备中等至高度自一致性,但跨模型一致性低。进一步分析显示,不同模型在援引道德原则方面呈现独特模式。这些发现凸显了在人工智能系统中实现一致道德推理的复杂性,并强调在医疗、陪伴等需伦理决策场景中,必须谨慎评估模型表现以规避潜在偏见与局限。
原文摘要 · Abstract (English)
The rapid adoption of large language models (LLMs) has spurred extensive research into their encoded moral norms and decision-making processes. Much of this research relies on prompting LLMs with survey-style questions to assess how well models are aligned with certain demographic groups, moral beliefs, or political ideologies. While informative, the adherence of these approaches to relatively superficial constructs tends to oversimplify the complexity and nuance underlying everyday moral dilemmas. We argue that auditing LLMs along more detailed axes of human interaction is of paramount importance to better assess the degree to which they may impact human beliefs and actions. To this end, we evaluate LLMs on complex, everyday moral dilemmas sourced from the "Am I the Asshole" (AITA) community on Reddit, where users seek moral judgments on everyday conflicts from other community members. We prompted seven LLMs to assign blame and provide explanations for over 10,000 AITA moral dilemmas. We then compared the LLMs' judgments and explanations to those of Redditors and to each other, aiming to uncover patterns in their moral reasoning. Our results demonstrate that large language models exhibit distinct patterns of moral judgment, varying substantially from human evaluations on the AITA subreddit. LLMs demonstrate moderate to high self-consistency but low inter-model agreement. Further analysis of model explanations reveals distinct patterns in how models invoke various moral principles. These findings highlight the complexity of implementing consistent moral reasoning in artificial systems and the need for careful evaluation of how different models approach ethical judgment. As LLMs continue to be used in roles requiring ethical decision-making such as therapists and companions, careful evaluation is crucial to mitigate potential biases and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。