arXiv:2608.03562cs.LG2026-08

让强化学习在实用函数出错时仍稳定,提升实际部署可靠性。

Robust General Utility for Reinforcement Learning

论文配图:Robust General Utility for Reinforcement Learning
图 1 · 摘自论文原文
  • 设计对抗性训练框架,应对实用函数设定偏差问题
  • 提出两种收敛算法,适配凹与非凹场景
  • 适用于大模型安全对齐、探索最大化等任务

强化学习中的广义效用框架通过优化策略诱导的占用测度的任意效用函数,拓展了经典强化学习的应用范围。然而,以往工作通常假设评估效用固定且正确指定。实践中,部署时使用的效用可能偏离训练阶段,造成鲁棒性缺口,而现有方法未予解决。为此,本文提出鲁棒广义效用强化学习,一种最小最大学习框架,使策略在预设的不确定性集中抵抗效用误设。该框架严格推广了标准广义效用强化学习,并通过合理选择效用不确定性集,统一视图包括奖励鲁棒强化学习和约束强化学习等多种现有框架。进一步,针对两种情形开发了可证明收敛的随机算法:对于凹效用,采用投影随机梯度下降-上升法,建立平稳性保证;对于更具挑战性的非凹情形,提出随机近似外梯度算法,缓解非凹性引发的病态行为,实现近似一阶平稳性的收敛。在大模型安全对齐与探索最大化任务上的实验验证了理论预测的收敛行为。

原文摘要 · Abstract (English)

Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications. However, previous work on general utility RL typically assumes the evaluation utility is fixed and correctly specified. In practice, the utility used at deployment can deviate from the training one, creating a robustness gap that prior work does not address. Motivated by this, we propose robust general-utility RL, a minimax learning framework that trains policies against utility misspecification within a prescribed uncertainty set. Our framework strictly generalizes standard general-utility RL while also providing a unified view of many existing RL frameworks, including reward-robust RL and constrained RL, through appropriate choices of the utility uncertainty set. We further develop provably convergent stochastic algorithms for two regimes. For concave utilities, we develop a projected stochastic gradient descent-ascent method and establish stationarity guarantees. For the more challenging nonconcave regime, we propose a stochastic prox-extragradient algorithm that mitigates ill-posed behavior induced by nonconcavity, with convergence guarantees to approximate first-order stationarity. Experiments on LLM safety alignment and exploration maximization tasks further corroborate the convergence behavior consistent with our theory.

强化学习鲁棒性大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。