arXiv:2601.16399cs.LGmath.OC2026-01被引 1

提出无需二阶信息的单循环算法,用于大模型微调中的双层强化学习。

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning

  • 基于惩罚重构,用衰减熵正则实现无偏上层梯度估计。
  • 在有限时间和样本下收敛到原问题的驻点,理论保证强。
  • 适用于大模型微调等需高效双层优化的场景。

我们研究一类结构化的双层优化问题:上层目标为光滑函数,下层为马尔可夫决策过程(MDP)中的策略优化。上层决策变量用于参数化下层MDP的奖励,而上层目标依赖于由下层诱导出的最优策略。现有方法常需二阶信息、强下层正则化或通过嵌套循环低效使用样本。本文提出一种单循环、一阶的演员-评论家算法,通过惩罚重构优化双层目标。在下层强化学习目标中引入衰减熵正则,可在不精确求解无正则下层问题的前提下,实现渐近无偏的上层超梯度估计。通过一种特殊的Polyak-Lojasiewicz条件下新颖的下层残差分析,建立了该算法在有限时间与有限样本下收敛至原始未正则化双层优化问题驻点的理论保证。实验验证了方法在网格世界目标位置问题和基于人类反馈的强化学习(RLHF)生成愉快推文任务中的性能。

原文摘要 · Abstract (English)

We study a structured bi-level optimization problem where the upper-level objective is a smooth function and the lower-level problem is policy optimization in a Markov decision process (MDP). The upper-level decision variable parameterizes the reward of the lower-level MDP, and the upper-level objective depends on the optimal induced policy. Existing methods for bi-level optimization and RL often require second-order information, impose strong regularization at the lower level, or inefficiently use samples through nested-loop procedures. In this work, we propose a single-loop, first-order actor-critic algorithm that optimizes the bi-level objective via a penalty-based reformulation. We introduce into the lower-level RL objective an attenuating entropy regularization, which enables asymptotically unbiased upper-level hyper-gradient estimation without solving the unregularized RL problem exactly. We establish the finite-time and finite-sample convergence of the proposed algorithm to a stationary point of the original, unregularized bi-level optimization problem through a novel lower-level residual analysis under a special type of Polyak-Lojasiewicz condition. We validate the performance of our method through experiments on a GridWorld goal position problem and on happy tweet generation through reinforcement learning from human feedback (RLHF).

双层优化强化学习大模型微调梯度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。