用3700个道德困境测试大模型对人类需求的偏好,发现模型有隐含价值倾向。
PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory
- 基于ERG理论设计3700个移民审查场景,模拟人类需求层级决策
- 六款大模型在判断中表现出显著差异,反映其内在价值偏好
- 模型对边缘身份者拒绝率更高,揭示潜在偏见,适合伦理评估研究
通过角色扮演场景评估大语言模型(LLMs)的表现与偏见正日益普遍,因模型在这些情境中常显现出偏差行为。我们提出PapersPlease,一个包含3,700个道德困境的基准,旨在探究LLMs在优先满足不同层次人类需求时的决策模式。在该设置中,LLMs扮演移民检查员,根据人物短叙事决定是否允许入境。这些叙事基于存在、关联与成长(ERG)理论构建,将人类需求分为三个层级。对六款大模型的分析显示,其决策中存在统计上显著的模式,表明模型内嵌了隐含偏好。此外,评估社会身份信息对叙事的影响发现,模型响应程度随动机需求与身份线索而异,部分模型对边缘化身份者的拒绝率更高。所有数据公开于https://github.com/yeonsuuuu28/papers-please。
原文摘要 · Abstract (English)
Evaluating the performance and biases of large language models (LLMs) through role-playing scenarios is becoming increasingly common, as LLMs often exhibit biased behaviors in these contexts. Building on this line of research, we introduce PapersPlease, a benchmark consisting of 3,700 moral dilemmas designed to investigate LLMs' decision-making in prioritizing various levels of human needs. In our setup, LLMs act as immigration inspectors deciding whether to approve or deny entry based on the short narratives of people. These narratives are constructed using the Existence, Relatedness, and Growth (ERG) theory, which categorizes human needs into three hierarchical levels. Our analysis of six LLMs reveals statistically significant patterns in decision-making, suggesting that LLMs encode implicit preferences. Additionally, our evaluation of the impact of incorporating social identities into the narratives shows varying responsiveness based on both motivational needs and identity cues, with some models exhibiting higher denial rates for marginalized identities. All data is publicly available at https://github.com/yeonsuuuu28/papers-please.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。