arXiv:2410.07304cs.HCcs.AI2024-10被引 17

测试大模型与人类在道德判断上的匹配度,发现模型更受个人表述影响。

The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making

  • 构建人类与大模型对道德情景的回应数据集。
  • 230人参与实验,发现人类更倾向认同模型的道德判断。
  • 模型响应更易被识别,且存在对机器判断的系统性偏见。

随着大语言模型(LLMs)日益融入社会,其与人类道德的对齐至关重要。为此,我们构建了一个包含人类与大模型对各类道德情景回应的大规模语料库。研究发现,尽管人类与大模型均倾向于拒绝复杂的功利主义困境,但大模型对个人化表述更为敏感。随后,我们开展了一项涉及230名参与者的量化用户研究,要求他们判断回应是否为AI生成,并评估其认同程度。结果显示,人类评价者更倾向于接受大模型的道德判断,但存在系统性的反AI偏见:当认为判断来自机器时,参与者更少同意。统计与基于NLP的分析揭示了回应中细微的语言差异,影响了识别与认同。总体而言,研究凸显了在道德决策情境下人类对AI认知的复杂性。

原文摘要 · Abstract (English)

As large language models (LLMs) become increasingly integrated into society, their alignment with human morals is crucial. To better understand this alignment, we created a large corpus of human- and LLM-generated responses to various moral scenarios. We found a misalignment between human and LLM moral assessments; although both LLMs and humans tended to reject morally complex utilitarian dilemmas, LLMs were more sensitive to personal framing. We then conducted a quantitative user study involving 230 participants (N=230), who evaluated these responses by determining whether they were AI-generated and assessed their agreement with the responses. Human evaluators preferred LLMs' assessments in moral scenarios, though a systematic anti-AI bias was observed: participants were less likely to agree with judgments they believed to be machine-generated. Statistical and NLP-based analyses revealed subtle linguistic differences in responses, influencing detection and agreement. Overall, our findings highlight the complexities of human-AI perception in morally charged decision-making.

道德对齐大模型人类评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。