arXiv:2504.10886cs.CYcs.AI2025-04中稿 · ICLR被引 13

研究大模型在道德困境中如何随人物角色变化决策,发现其行为易受政治立场影响。

Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment

  • 通过不同社会人口角色模拟测试大模型道德决策模式
  • 模型在关键道德任务中决策偏差显著大于人类
  • 政治身份主导模型决策方向,提示部署风险

将具备自主性的大语言模型(LLMs)应用于现实场景时,其行为是否符合人类道德判断成为关键问题。本研究在道德机器实验的不同情境下,考察了反映多元社会人口特征的人物角色对LLM决策的影响。结果显示,不同人物角色下,大模型的道德判断存在显著差异,且在关键任务中的决策波动幅度远超人类。数据还揭示一种政党分化现象:政治身份主导了模型决策的方向与程度。这表明,当模型被用于涉及道德抉择的应用时,可能放大社会偏见并带来伦理风险。

原文摘要 · Abstract (English)

Deploying large language models (LLMs) with agency in real-world applications raises critical questions about how these models will behave. In particular, how will their decisions align with humans when faced with moral dilemmas? This study examines the alignment between LLM-driven decisions and human judgment in various contexts of the moral machine experiment, including personas reflecting different sociodemographics. We find that the moral decisions of LLMs vary substantially by persona, showing greater shifts in moral decisions for critical tasks than humans. Our data also indicate an interesting partisan sorting phenomenon, where political persona predominates the direction and degree of LLM decisions. We discuss the ethical implications and risks associated with deploying these models in applications that involve moral decisions.

大模型对齐道德推理角色建模伦理风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。