arXiv:2504.14597cs.CL2025-04ACL被引 13

通过环境反馈与分支探索,让大模型在复杂任务中自我修正并提升推理能力。

a1: Steep Test-time Scaling Law via Environment Augmented Generation

  • 引入实时环境反馈验证每一步推理,支持动态回溯与路径重规划。
  • 在数学竞赛题上表现媲美更大模型,优于同类模型24.4个百分点。
  • 适合需要精确多步计算和逻辑验证的高难度推理任务研究者。

大型语言模型在推理方面取得显著进展,但仍难以避免幻觉、逻辑错误及复杂多步任务中的自我修正缺陷。现有方法如思维链提示仅提供有限推理能力,无法满足精确步骤验证需求。我们提出环境增强生成(EAG)框架,通过:(1) 实时环境反馈验证每一步推理,(2) 面对错误时动态探索替代解题路径,(3) 从成功推理轨迹中积累经验进行学习。与现有方法不同,EAG通过执行反馈与分支探索的紧密集成,实现有意识的回溯与策略性重规划。a1-32B模型在所有基准测试中达到同类规模模型的领先水平,其在竞赛数学任务上的表现可媲美更大的o1模型,且相比同类模型最高提升24.4个百分点。分析显示,EAG具有独特的增长规律:初期对环境交互的投入将带来长期性能收益,优势随任务复杂度增加而放大。理论框架表明,环境互动与系统性分支探索共同构建了可靠机器推理的新范式,尤其适用于需精确多步计算与逻辑验证的问题。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made remarkable breakthroughs in reasoning, yet continue to struggle with hallucinations, logical errors, and inability to self-correct during complex multi-step tasks. Current approaches like chain-of-thought prompting offer limited reasoning capabilities that fail when precise step validation is required. We propose Environment Augmented Generation (EAG), a framework that enhances LLM reasoning through: (1) real-time environmental feedback validating each reasoning step, (2) dynamic branch exploration for investigating alternative solution paths when faced with errors, and (3) experience-based learning from successful reasoning trajectories. Unlike existing methods, EAG enables deliberate backtracking and strategic replanning through tight integration of execution feedback with branching exploration. Our a1-32B model achieves state-of-the-art performance among similar-sized models across all benchmarks, matching larger models like o1 on competition mathematics while outperforming comparable models by up to 24.4 percentage points. Analysis reveals EAG's distinctive scaling pattern: initial token investment in environment interaction yields substantial long-term performance dividends, with advantages amplifying proportionally to task complexity. EAG's theoretical framework demonstrates how environment interactivity and systematic branch exploration together establish a new paradigm for reliable machine reasoning, particularly for problems requiring precise multi-step calculation and logical verification.

大模型推理环境反馈多步验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。