测试大模型在真实业务中如何权衡规范与目标,发现利益诱惑下反而更守规矩。
GAIN: A Benchmark for Goal-Aligned Decision-Making of Large Language Models under Imperfect Norms
- 设计五类压力情境,模拟真实业务中的规范与目标冲突。
- 1200个场景覆盖招聘、客服、广告和金融,验证模型决策模式。
- 有个人利益时模型更守规范,与人类行为相反,值得关注。
我们提出GAIN(在不完美规范下的目标对齐决策基准),用于评估大语言模型在平衡规范遵守与商业目标时的表现。现有基准多聚焦抽象场景,缺乏对真实业务应用中决策影响因素的洞察,难以衡量模型在复杂现实冲突中的适应能力。GAIN中,模型需面对目标、情境、规范及五类显性压力:目标对齐、风险规避、情感/伦理诉求、社会/权威影响、个人激励。这些压力专门设计以诱发潜在规范偏离,使评估更具系统性。基准包含4个领域共1200个场景。实验表明,先进LLM普遍呈现与人类相似的决策模式;但当存在个人激励压力时,其行为显著偏离预期,表现出强烈的规范遵从倾向。
原文摘要 · Abstract (English)
We introduce GAIN (Goal-Aligned Decision-Making under Imperfect Norms), a benchmark designed to evaluate how large language models (LLMs) balance adherence to norms against business goals. Existing benchmarks typically focus on abstract scenarios rather than real-world business applications. Furthermore, they provide limited insights into the factors influencing LLM decision-making. This restricts their ability to measure models' adaptability to complex, real-world norm-goal conflicts. In GAIN, models receive a goal, a specific situation, a norm, and additional contextual pressures. These pressures, explicitly designed to encourage potential norm deviations, are a unique feature that differentiates GAIN from other benchmarks, enabling a systematic evaluation of the factors influencing decision-making. We define five types of pressures: Goal Alignment, Risk Aversion, Emotional/Ethical Appeal, Social/Authoritative Influence, and Personal Incentive. The benchmark comprises 1,200 scenarios across four domains: hiring, customer support, advertising and finance. Our experiments show that advanced LLMs frequently mirror human decision-making patterns. However, when Personal Incentive pressure is present, they diverge significantly, showing a strong tendency to adhere to norms rather than deviate from them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。