arXiv:2509.25369cs.CLcs.AI2025-09中稿 · ICLR被引 13

通过自动构造价值冲突场景,揭示大模型在不同情境下的价值优先级。

Generative Value Conflicts Reveal LLM Priorities

  • 构建自动化管道ConflictScope,生成模型需在价值间权衡的对话场景。
  • 开放式问答中模型更倾向用户自主而非安全保护,倾向性显著变化。
  • 在系统提示中加入价值排序可提升14%对齐度,有助于改善冲突应对。

以往研究致力于将大语言模型助手与特定价值观对齐,但实际部署中常面临价值冲突。由于现有对齐数据集缺乏价值冲突样本,我们提出ConflictScope——一种自动评估大模型在不同价值间如何优先排序的流水线。给定一组用户定义的价值,ConflictScope自动生成模型需在其中两个价值间做权衡的情境,并以大模型撰写的“用户提问”形式向目标模型发起挑战,通过分析其自由文本回答,推断其对价值集合的偏好排序。对比多选与开放式评估发现,在更开放的冲突设置下,模型倾向于支持用户自主等个人价值,而非伤害避免等保护性价值。然而,在系统提示中加入明确的价值排序可使模型行为与目标排序对齐度提升14%,表明系统提示可在价值冲突中实现中等程度的对齐。本工作强调了评估模型价值优先级的重要性,并为该方向的研究奠定了基础。

原文摘要 · Abstract (English)

Past work seeks to align large language model (LLM)-based assistants with a target set of values, but such assistants are frequently forced to make tradeoffs between values when deployed. In response to the scarcity of value conflict in existing alignment datasets, we introduce ConflictScope, an automatic pipeline to evaluate how LLMs prioritize different values. Given a user-defined value set, ConflictScope automatically generates scenarios in which a language model faces a conflict between two values sampled from the set. It then prompts target models with an LLM-written "user prompt" and evaluates their free-text responses to elicit a ranking over values in the value set. Comparing results between multiple-choice and open-ended evaluations, we find that models shift away from supporting protective values, such as harmlessness, and toward supporting personal values, such as user autonomy, in more open-ended value conflict settings. However, including detailed value orderings in models' system prompts improves alignment with a target ranking by 14%, showing that system prompting can achieve moderate success at aligning LLM behavior under value conflict. Our work demonstrates the importance of evaluating value prioritization in models and provides a foundation for future work in this area.

大模型对齐价值冲突提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。