arXiv:2604.01161cs.LG2026-04被引 4

上下文会悄悄压缩大模型的推理过程,影响其可靠性。

Reasoning Shift: How Context Silently Shortens LLM Reasoning

  • 在不同上下文中,模型推理长度最多缩短65%
  • 上下文干扰导致自我验证行为减少,风险上升
  • 适合关注模型鲁棒性与上下文管理的研究者

大型语言模型(LLMs)在复杂长时推理任务中表现出测试时扩展能力,如延长推理链条和自我验证。然而,这些推理行为的鲁棒性尚未充分研究。我们系统评估了多种推理模型在三种场景下的表现:(1) 包含冗长无关上下文的问题;(2) 多轮对话中独立任务;(3) 复杂任务中的子任务。结果发现,在不同上下文条件下,相同问题的推理链长度相比孤立呈现时最多缩短65%。细粒度分析显示,这种压缩与自我验证和不确定性管理行为(如双重检查)下降相关。虽然对简单问题性能无损,但在挑战性任务中可能带来风险。此外,针对性的监督微调可部分缓解无关上下文的负面影响。本研究呼吁重视推理模型的鲁棒性及上下文管理问题。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks. However, the robustness of these reasoning behaviors remains underexplored. To investigate this, we conduct a systematic evaluation of multiple reasoning models across three scenarios: (1) problems augmented with lengthy, irrelevant context; (2) multi-turn conversational settings with independent tasks; and (3) problems presented as a subtask within a complex task. We observe an interesting phenomenon: reasoning models tend to produce much shorter reasoning traces (up to 65%) for the same problem under different context conditions compared to the traces produced when the problem is presented in isolation. A finer-grained analysis reveals that this compression is associated with a decrease in self-verification and uncertainty management behaviors, such as double-checking. While this behavioral shift does not compromise performance on straightforward problems, it might affect performance on more challenging tasks. Additionally, we show that targeted supervised fine-tuning partially mitigates the adverse effects of irrelevant context. We hope our findings draw additional attention to both the robustness of reasoning models and the problem of context management for LLMs and LLM-based agents.

大模型推理上下文影响模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。