arXiv:2601.07226cs.AIcs.CL2026-01被引 10

顶尖推理模型在噪声上下文下表现骤降,新基准揭示其脆弱性。

Lost in the Noise: How Reasoning Models Fail with Contextual Distractors

  • 构建多类型噪声基准NoisyBench,评估模型在11个任务中的鲁棒性。
  • 顶级模型在干扰项下性能最高下降80%,且代理流程会放大错误。
  • 提出RARE奖励机制,通过识别有用信息提升抗噪能力,适合构建可靠智能体。

近期推理模型与智能体系统日益依赖外部信息,但当前输入常含固有噪声,现有清洁基准无法反映这一现实。我们提出NoisyBench,一个涵盖11个数据集的综合基准,系统评估模型在RAG、推理、对齐和工具使用任务中面对随机文档、无关对话历史及硬负样本等噪声类型时的鲁棒性。评估发现,顶尖模型在上下文干扰下性能最高下降80%。关键的是,代理工作流常因过度信任噪声工具输出而放大错误,干扰项甚至可引发非对抗性误对齐。提示工程、上下文设计、SFT及结果奖励强化学习均无法保障鲁棒性;相比之下,我们提出的理性感知奖励(RARE)通过激励模型识别有用信息显著增强抗噪能力。最后,我们发现测试时计算量越大,噪声环境下性能反而越差,并通过注意力可视化证实模型过度关注干扰词,为构建下一代稳健推理智能体提供关键洞见。

原文摘要 · Abstract (English)

Recent advances in reasoning models and agentic AI systems have led to an increased reliance on diverse external information. However, this shift introduces input contexts that are inherently noisy, a reality that current sanitized benchmarks fail to capture. We introduce NoisyBench, a comprehensive benchmark that systematically evaluates model robustness across 11 datasets in RAG, reasoning, alignment, and tool-use tasks against diverse noise types, including random documents, irrelevant chat histories, and hard negative distractors. Our evaluation reveals a catastrophic performance drop of up to 80% in state-of-the-art models when faced with contextual distractors. Crucially, we find that agentic workflows often amplify these errors by over-trusting noisy tool outputs, and distractors can trigger emergent misalignment even without adversarial intent. We find that prompting, context engineering, SFT, and outcome-reward only RL fail to ensure robustness; in contrast, our proposed Rationale-Aware Reward (RARE) significantly strengthens resilience by incentivizing the identification of helpful information within noise. Finally, we uncover an inverse scaling trend where increased test-time computation leads to worse performance in noisy settings and demonstrate via attention visualization that models disproportionately focus on distractor tokens, providing vital insights for building the next generation of robust, reasoning-capable agents.

推理模型噪声鲁棒性智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。