arXiv:2601.21439cs.AI2026-01被引 1

LLM在规则决策中对情绪干扰有惊人鲁棒性,与表面脆弱形成悖论。

The Paradox of Robustness: Decoupling Rule-Based Logic from Affective Noise in High-Stakes Decision-Making

  • 用三类高风险场景测试提示扰动,发现模型对情感框架影响不敏感
  • 效应量极小(Cohen's h=0.003),仅为人类偏差的1/100
  • 适合关注AI决策可靠性或政策评估的研究者参考

尽管大语言模型(LLMs)被广泛记录为对微小提示扰动敏感且易出现讨好式对齐,其在重要、规则驱动的决策中的鲁棒性仍缺乏研究。我们揭示了一个显著的「鲁棒性悖论」:尽管存在词汇上的脆弱性,对齐后的LLMs在规则驱动的制度决策中表现出对情感框架效应的强大鲁棒性。在医疗、金融和教育三个高风险领域,采用受控扰动框架进行测试,发现效应量极小(Cohen's h = 0.003),相较于人类类似情境中的显著偏差(h 在 [0.3, 0.8] 之间)约小两个数量级。该不变性在八种不同训练范式的模型中均持续存在,表明导致讨好行为和提示敏感性的机制并未传递到逻辑约束满足的失败上。为探查此发现边界,额外开展两项评审驱动研究:五场景移民扩展显示+0.8个百分点的小但统计显著偏移,仍在预设的±3百分点实际等效区域(ROPE)内;筛查级对抗叙事试点则未发现更强生成提示下的决策显著变化。研究发布核心基准(9个基础场景 × 18种条件变体 = 162个唯一提示)、代码与数据,以支持可复现评估。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) are widely documented to be sensitive to minor prompt perturbations and prone to sycophantic alignment, their robustness in consequential, rule-bound decision-making remains under-explored. We uncover a striking "Paradox of Robustness": despite their known lexical brittleness, aligned LLMs exhibit strong robustness to emotional framing effects in rule-bound institutional decision-making. Using a controlled perturbation framework across three high-stakes domains (healthcare, finance, and education), we find a negligible effect size (Cohen's h = 0.003) compared to the substantial biases observed in analogous human contexts (h in [0.3, 0.8]), approximately two orders of magnitude smaller. This invariance persists across eight models with diverse training paradigms, suggesting the mechanisms driving sycophancy and prompt sensitivity do not translate to failures in logical constraint satisfaction. While LLMs may be "brittle" to how a query is formatted, they appear considerably more stable against affective attempts to bias rule-bound decisions. To probe the boundary of this finding, we add two reviewer-driven side studies. A five-scenario immigration extension yields a small but statistically detectable +0.8 percentage point shift that remains within a pre-specified +/-3 percentage point Region of Practical Equivalence (ROPE), while a screening-level adversarial narrative pilot finds no meaningful decision shift under stronger LLM-generated prompts. We release a core benchmark (9 base scenarios x 18 condition variants = 162 unique prompts), code, and data to facilitate replicable evaluation.

决策系统鲁棒性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。