LLM决策易受表述方式影响,新方法通过稳定价值导向降低偏差。
Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

- 在三种语义框架下构建大规模测试集,发现模型决策翻转率高达28.6%
- 传统提示和激活层干预反而加剧敏感性,效果适得其反
- 提出代表层方法Valign,通过锚定价值先验抑制框架干扰
大型语言模型(LLMs)在法律推理等高风险决策场景中应用日益广泛,但事实相同仅表述不同的输入会显著动摇模型决策。为系统研究此问题,我们提出Fragile——一个大规模基准,控制性地分离出三种语义框架维度:价值染色叙述、时间切片与叙事生动性。实验显示,LLMs对框架高度敏感,平均决策翻转率达28.6%。简单采用提示层或激活层干预不仅无效,甚至加剧敏感性。为此,我们提出Valign,一种表示层方法,通过锚定稳定的值先验,引导隐藏状态向模型自身价值一致方向演化,并从隐藏状态中投影出时间-生动性敏感方向。该方法持续降低由框架引发的决策翻转,表明鲁棒缓解需直接干预内部作用路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factually equivalent inputs is critical. However, we find that fact-preserved but differently framed inputs can significantly destabilize LLM decisions. To systematically investigate this problem, we introduce Fragile, a large-scale benchmark that isolates fact-preserving semantic framing across three controlled dimensions: value-tinted narration, temporal slice, and narrative vividness. Our experiments reveal a high susceptibility of LLMs to framing, with an average decision flip rate of 28.6%. We find that simple prior prompt-level and activation-level interventions not only fail to suppress framing sensitivity but actively amplify it. We therefore propose Valign, a representation-level method that explicitly targets these framing dimensions by anchoring decisions to a stable value prior, steering hidden states toward the model's value-consistent direction, and projecting out temporal-vividness-sensitive directions from the model's hidden states. Valign consistently reduces framing-induced decision flips, demonstrating that robust mitigation requires directly targeting the internal pathways in which framing operates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。