arXiv:2605.04065cs.CLcs.ET2026-05ACL

让大模型自我进化时,能动态调整奖励和学习信号。

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs

论文配图:Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
图 1 · 摘自论文原文
  • 用自由能原理自适应调节奖励,平衡探索与共识。
  • 根据采样奖励的统计特征动态调整学习信号。
  • 在数学推理任务上比其他方法高0.5至3.5分。

无监督强化学习已成为推动大语言模型自我改进的有前景范式。然而,现有方法在训练过程中难以适应模型推理能力的变化,缺乏真实标签监督时易导致策略优化偏离正轨。为此,我们提出FREIA,一种基于两项关键创新的新型强化学习算法:(1) 自由能驱动奖励(FER)依据自由能原理,自适应地平衡一致性与探索性;(2) 自适应优势塑造(AAS)根据采样奖励的统计特性动态调整学习信号。在九个数据集、三个推理任务上的实证评估表明,FREIA优于其他无监督强化学习基线方法。尤其在数学推理任务中,使用DeepSeek-R1-Distill-Qwen-1.5B模型时,其Pass@1平均得分比其他方法高出0.5至3.5分。

原文摘要 · Abstract (English)

Unsupervised reinforcement learning (RL) has emerged as a promising paradigm for enabling self-improvement in large language models (LLMs). However, existing unsupervised RL-based methods often lack the capacity to adapt to the model's evolving reasoning capabilities during training. Therefore, these methods can misdirect policy optimization in the absence of ground-truth supervision. To address this issue, we introduce FREIA, a novel RL-based algorithm built on two key innovations: (1) Free Energy-Driven Reward (FER) adapts rewards to balance consensus and exploration based on the Free Energy Principle. (2) Adaptive Advantage Shaping (AAS) adaptively adjusts learning signals based on the statistical characteristics of sampled rewards. Empirical evaluations on nine datasets across three reasoning tasks showcase that FREIA outperforms other unsupervised RL-based baselines. Notably, in mathematical reasoning tasks, FREIA surpasses other methods by an average of 0.5 to 3.5 points in Pass@1 using the DeepSeek-R1-Distill-Qwen-1.5B model.

强化学习大模型推理自适应奖励无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。