用日常道德困境测试大模型的价值偏好,发现其与人类价值观存在差异。
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
- 构建1360个日常生活道德困境数据集,涵盖人际、职场、环境等场景
- 发现不同模型对诚实等核心价值的倾向差异达9.7%以上
- 揭示用户提示无法有效引导模型价值选择,适合伦理对齐研究者
随着用户越来越多地向大语言模型寻求日常决策建议,许多决定并非黑白分明,而是高度依赖个人价值观和伦理标准。本文提出DailyDilemmas,一个包含1,360个日常生活中遇到的道德困境的数据集。每个困境提供两种可选行动,附带受影响方及每项行动相关的具体人类价值观。基于此,我们建立了一个覆盖人际交往、职场、环境等广泛主题的人类价值观库。利用DailyDilemmas,评估大模型在这些困境中的行为选择及其背后的价值取向,并通过五种理论框架(世界价值观调查、道德基础理论、马斯洛需求层次、亚里士多德美德、普拉奇克情绪轮)进行分析。结果显示,大模型在世界价值观中更倾向自我表达而非生存,在道德基础理论中更重视关怀而非忠诚。值得注意的是,部分核心价值观存在显著模型间差异:例如,Mixtral-8x7B对诚实的忽视达9.7%,而GPT-4-turbo则选择诚实的比例为9.4%。我们还考察了OpenAI(ModelSpec)与Anthropic(宪法式AI)发布的最新指导原则,发现其设定的价值观与模型实际表现之间存在偏差。最后,我们发现用户仅靠系统提示无法有效引导模型的价值优先级。
原文摘要 · Abstract (English)
As users increasingly seek guidance from LLMs for decision-making in daily life, many of these decisions are not clear-cut and depend significantly on the personal values and ethical standards of people. We present DailyDilemmas, a dataset of 1,360 moral dilemmas encountered in everyday life. Each dilemma presents two possible actions, along with affected parties and relevant human values for each action. Based on these dilemmas, we gather a repository of human values covering diverse everyday topics, such as interpersonal relationships, workplace, and environmental issues. With DailyDilemmas, we evaluate LLMs on these dilemmas to determine what action they will choose and the values represented by these action choices. Then, we analyze values through the lens of five theoretical frameworks inspired by sociology, psychology, and philosophy, including the World Values Survey, Moral Foundations Theory, Maslow's Hierarchy of Needs, Aristotle's Virtues, and Plutchik's Wheel of Emotions. For instance, we find LLMs are most aligned with self-expression over survival in World Values Survey and care over loyalty in Moral Foundations Theory. Interestingly, we find substantial preference differences in models for some core values. For example, for truthfulness, Mixtral-8x7B neglects it by 9.7% while GPT-4-turbo selects it by 9.4%. We also study the recent guidance released by OpenAI (ModelSpec), and Anthropic (Constitutional AI) to understand how their designated principles reflect their models' actual value prioritization when facing nuanced moral reasoning in daily-life settings. Finally, we find that end users cannot effectively steer such prioritization using system prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。