arXiv:2507.07017cs.AI2025-07被引 41

通过精准探索提升大模型推理稳定性

First Return, Entropy-Eliciting Explore

  • 识别推理路径中的高不确定性点,针对性展开试错
  • 在AIME24上使正确推理轨迹比例显著提升
  • 无需密集标注,适合需要可靠推理的场景

基于可验证奖励的强化学习(RLVR)虽能增强大语言模型(LLM)的推理能力,但面临探索不稳定的挑战。本文提出FR3E(首次返回、熵激励探索)框架,通过识别推理轨迹中高不确定性的决策点,进行有针对性的模拟回溯,构建语义明确的中间反馈。该方法无需依赖密集监督,即可提供精准引导。在数学推理基准(AIME24)上的实验表明,FR3E能实现更稳定训练,生成更长且连贯的响应,并显著提高完全正确推理路径的比例。结果证明该框架可通过更稳健、结构化的探索有效提升大模型推理能力。

原文摘要 · Abstract (English)

Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (First Return, Entropy-Eliciting Explore), a structured exploration framework that identifies high-uncertainty decision points in reasoning trajectories and performs targeted rollouts to construct semantically grounded intermediate feedback. Our method provides targeted guidance without relying on dense supervision. Empirical results on mathematical reasoning benchmarks(AIME24) show that FR3E promotes more stable training, produces longer and more coherent responses, and increases the proportion of fully correct trajectories. These results highlight the framework's effectiveness in improving LLM reasoning through more robust and structured exploration.

推理增强强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。