发现生成模型中的轨迹平衡与信任PCL等价,统一了强化学习优化理论。
Relative Trajectory Balance is equivalent to Trust-PCL
- 通过理论推导证明轨迹平衡与信任PCL在KL正则化下本质相同。
- 在示例任务中验证两者性能相当,支持新视角的合理性。
- 适合研究生成模型与强化学习交叉的学者参考。
近期生成建模进展表明,强化学习(RL)在微调阶段至关重要,尤其是带有KL正则化的策略在自回归与扩散模型中均表现优异。与此同时,相对轨迹平衡(RTB)作为生成流网络(GFlowNets)中用于提升序列生成模型微调效果的新目标被提出。本文基于先前将GFlowNets与最大熵强化学习关联的工作,首次建立RTB与一种具有KL正则化的离线策略强化学习方法——信任PCL之间的等价关系。该等价性将RTB置于更广泛的KL正则化强化学习理论框架中,并厘清了其与早期方法的关系。借助这一洞察,我们重新审视了原RTB论文中的一个典型例子,发现KL正则化强化学习方法可达到相似性能,为此前结论提供了替代解释。
原文摘要 · Abstract (English)
Recent progress in generative modeling has highlighted the importance of Reinforcement Learning (RL) for fine-tuning, with KL-regularized methods in particular proving to be highly effective for both autoregressive and diffusion models. Complementing this line of work, the Relative Trajectory Balance (RTB) objective was recently introduced in the context of Generative Flow Networks (GFlowNets) to serve the same role of improving fine-tuning in sequential generative models. Building on prior work linking GFlowNets and maximum-entropy RL, we establish in this paper an equivalence between RTB and Trust-PCL, an off-policy RL method with KL regularization. This equivalence situates RTB within the broader theoretical landscape of KL-regularized RL, and clarifies its relationship to earlier methods. Leveraging this insight, we revisit an illustrative example from the RTB paper and show that KL-regularized RL methods achieve comparable performance, offering an alternative perspective to what was previously reported.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。