arXiv:2607.12963cs.CL2026-07被引 2

看似鲁棒的模型其实对无关上下文极度敏感,局部预测会翻转。

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

论文配图:The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
图 1 · 摘自论文原文
  • 用随机字符伪词测试模型,发现上下文扰动引发局部预测变化。
  • 整体准确率不变,但少数样本预测翻转,性能或升或降。
  • 适合关注模型可靠性与安全性的研究者和开发者。

随着大语言模型能力增强,其越来越多地部署于包含大量部分无关上下文的复杂场景中。在受控实验中,我们发现当前最先进模型在整体上似乎对任务无关上下文具有鲁棒性:在基准问题前添加上下文后,整体准确率几乎不变。然而,这种整体稳定性掩盖了个体样本的显著不稳定性。即使是语义无意义的伪词(由随机字符组合而成),也能显著改变少数样本的模型预测,导致部分样本性能下降,另一些则提升。这一双向影响在多种模型和数据集上均持续存在,但受影响样本具有高度模型特异性。我们进一步表明,这种不稳定性受上下文类型、长度、测试时计算量及模型发展阶段的影响。综上,我们的研究揭示了被整体准确率掩盖的上下文诱发尾部风险,呼吁开展针对每个样本的可靠性评估。

原文摘要 · Abstract (English)

As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. This aggregate stability, however, masks significant per-example instability. Even semantically meaningless pseudo-words, formed by randomly combining characters, can markedly shift model predictions on a small fraction of examples, degrading performance on some while improving it on others. This two-sided effect holds consistently across a wide range of models and datasets, yet the affected examples are largely model-specific. We further show that this instability is modulated by context type, context length, test-time compute, and model development stage. Together, our findings reveal context-induced tail risks concealed by aggregate accuracy, motivating per-example reliability evaluation of language models.

大模型上下文干扰可靠性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。