评测大模型在高风险困境中的多视角决策能力,发现其在价值冲突中表现有限。
CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives
- 构建包含345个高影响困境的CLASH数据集,涵盖3795种不同价值观视角。
- 顶级模型在复杂价值判断中准确率仅24%至51%,难以处理矛盾与价值转变。
- 第一人称视角更易引导安全等特定价值,但多数认知策略不适用于价值推理。
在涉及价值冲突的高风险领域,人类决策已具挑战性,对人工智能而言更是难题,但现有研究多局限于日常情境。为此,我们提出CLASH(基于角色视角的高风险情境大模型评估),一个精心构建的数据集,包含345个高影响力困境及3,795个来自不同价值观的个体视角。该数据集支持对价值决策中关键但未被充分研究的方面进行分析,包括决策矛盾、心理不适感,以及角色视角中价值随时间的变化。通过在14个非思考型与思考型模型上进行基准测试,我们发现:(1)即使强大多数模型如GPT-5和Claude-4-Sonnet,在处理矛盾决策时准确率也仅为24.06%和51.01%;(2)尽管能合理预测心理不适,但对价值转移的视角理解不足;(3)数学与策略类有效认知行为无法迁移至价值推理,反而出现早期承诺与过度承诺等新失败模式;(4)模型可引导性与其自身价值偏好显著相关;(5)从第三方视角推理时模型更易引导,但某些价值(如安全)在第一人称框架下表现更优。
原文摘要 · Abstract (English)
Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been limited to everyday scenarios. To close this gap, we introduce CLASH (Character perspective-based LLM Assessments in Situations with High-stakes), a meticulously curated dataset consisting of 345 high-impact dilemmas along with 3,795 individual perspectives of diverse values. CLASH enables the study of critical yet underexplored aspects of value-based decision-making processes, including understanding of decision ambivalence and psychological discomfort as well as capturing the temporal shifts of values in the perspectives of characters. By benchmarking 14 non-thinking and thinking models, we uncover several key findings. (1) Even strong proprietary models, such as GPT-5 and Claude-4-Sonnet, struggle with ambivalent decisions, achieving only 24.06 and 51.01 accuracy. (2) Although LLMs reasonably predict psychological discomfort, they do not adequately comprehend perspectives involving value shifts. (3) Cognitive behaviors that are effective in the math-solving and game strategy domains do not transfer to value reasoning. Instead, new failure patterns emerge, including early commitment and overcommitment. (4) The steerability of LLMs towards a given value is significantly correlated with their value preferences. (5) Finally, LLMs exhibit greater steerability when reasoning from a third-party perspective, although certain values (e.g., safety) benefit uniquely from first-person framing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。