arXiv:2505.17621cs.LG2025-05被引 34

用内在动机引导探索,让大模型推理更高效

Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration

  • 设计密度奖励机制,减少对熟悉路径的依赖
  • 在AIME 2024上提升22.23%推理准确率
  • 适合需要深度迭代推理的任务场景

强化学习已成为提升大语言模型推理能力的关键方法。然而,主流方法如近端策略优化和组相对策略优化存在奖励稀疏、依赖结果反馈、探索激励弱等问题,尤其在复杂任务中表现受限。本文提出IMAGINE框架,通过三项创新实现密集奖励与有效探索:轨迹感知的探索奖励可高效降低逐标记偏差;基于错误的奖励分配机制能提升困难样本的探索效率并稳定训练;优势保留的融合机制确保学习过程中的分布一致性。在四个公开数据集上的实验表明,IMAGINE在AIME 2024上性能提升22.23%。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has become a key approach for enhancing the reasoning capabilities of large language models. However, prevalent RL approaches like proximal policy optimization and group relative policy optimization suffer from sparse, outcome-based rewards and weak exploration incentives, limiting their effectiveness. Specifically, sparse rewards offer limited feedback, especially on difficult problems, and introduce biases favoring familiar trajectories over novel reasoning paths. These issues critically undermine performance on complex tasks that inherently require iterative reasoning. To overcome these challenges, we propose Intrinsic MotivAtion Guided exploratIoN for Enhanced reasoning (IMAGINE), which delivers dense rewards and encourages exploration. IMAGINE introduces three innovations: a trajectory-aware exploration reward that reduces token-level bias efficiently; an error-conditioned reward allocation that promotes efficient exploration on hard samples while stabilizing training; and an advantage-preserving integration mechanism that retains distributional integrity during learning. Experiments on four public datasets show that IMAGINE improves performance by 22.23% on AIME 2024.

大模型推理强化学习探索机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。