arXiv:2505.22290cs.AIcs.CL2025-05被引 2

通过搜索+推理放大,让难解推理题成功率提升30倍

Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling

  • 用上下文搜索引导模型逐步推理
  • 在硬核任务上成功率最高提升30倍
  • 适合研究模型极限与评估方法的学者

近期研究指出,即使经过长链推理训练,大语言模型(LLMs)在复杂推理任务上仍面临显著挑战。现有文献多依赖简单上下文学习进行评估,忽视了激发模型深度推理能力的先进方法,导致性能被低估。本文系统探索了上下文搜索与测试时扩展(test-time scaling)结合的潜力。结果表明,在引入内部缩放机制后,通过高级上下文搜索提示,可实现对以往被认为“不可解”任务的颠覆性突破(原成功率低于5%)。实证显示,在受控的NP-hard任务和真实世界规划基准上,该方法相比无外部机制的先前结果,成功率最高提升30倍;理论分析进一步证明,该组合显著拓展了可解推理问题的复杂度类别。研究挑战了当前对大模型能力的固有认知,呼吁重新审视评估范式,建立更全面的评测体系以准确揭示当代大模型的真实推理边界。

原文摘要 · Abstract (English)

Recent research has highlighted that Large Language Models (LLMs), even when trained to generate extended long reasoning steps, still face significant challenges on hard reasoning problems. However, much of the existing literature relies on direct prompting with simple in-context learning examples for evaluation, which largely overlooks advanced techniques to elicit LLMs' deliberate reasoning before drawing conclusions that LLMs hit a performance ceiling. In this paper, we systematically explore the combined potential of in-context search and test-time scaling on super hard reasoning tasks. We find that by employing advanced in-context search prompting to LLMs augmented with internal scaling, one can achieve transformative performance breakthroughs on tasks previously deemed "unsolvable" (e.g., reported success rates below 5%). We provide both empirical results and theoretical analysis of how this combination can unleash LLM reasoning capabilities: i) Empirically, on controlled NP-hard tasks and complex real-world planning benchmarks, our approach achieves up to a 30x improvement in success rates compared to previously reported results without any external mechanisms; ii) Theoretically, we show that in-context search prompting, when combined with internal scaling, significantly extends the complexity class of solvable reasoning problems. These findings challenge prevailing assumptions about the limitations of LLMs on complex tasks, indicating that current evaluation paradigms systematically underestimate their true potential. Our work calls for a critical reassessment of how LLM reasoning is benchmarked and a more robust evaluation strategy that fully captures the true capabilities of contemporary LLMs, which can lead to a better understanding of their operational reasoning boundaries in real-world deployments.

大模型推理测试时扩展上下文搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。