提升推荐与生成模型的强化学习安全性与效率
Safe, Efficient, and Robust Reinforcement Learning for Ranking and Diffusion Models
- 基于上下文老虎机框架,设计安全可信赖的排序算法
- 提出新评估方法,使生成结果更贴合文本描述,且样本效率高
- 适合关注模型安全、生成质量与训练效率的研究者
本论文研究如何设计安全、高效且鲁棒的强化学习方法。从上下文老虎机统一视角出发,解决排序推荐与文生图扩散模型两大应用问题。第一部分构建暴露相关的泛化界,提出反事实风险最小化目标,确保策略性能不低于记录策略,即使在反馈稀疏情况下也成立;该保障进一步扩展至双重稳健估计器,在对抗性或错误建模条件下仍能保证安全,并为从业者提供可控制的效用损失上限。第二部分聚焦单动作老虎机,将多种离策略估计算法统一于基线修正框架下,提出闭式最优基线,显著降低评估与策略梯度方差,提升离策略学习可靠性。第三部分系统分析生成式强化学习中效率与效果的权衡,基于PPO与REINFORCE的比较,提出留一法PPO(LOOP)算法,结合多条扩散轨迹与类REINFORCE基线,嵌入PPO裁剪目标中,实现与PPO相当的样本效率,同时生成结果更忠实于文本属性。
原文摘要 · Abstract (English)
This dissertation investigates how reinforcement learning (RL) methods can be designed to be safe, sample-efficient, and robust. Framed through the unifying perspective of contextual-bandit RL, the work addresses two major application domains - ranking and recommendation, and text-to-image diffusion models. The first part of the thesis develops theory and algorithms for safe deployment in ranking systems. An exposure-based generalisation bound is derived, leading to a counterfactual risk-minimisation objective whose solution is guaranteed not to underperform the logging policy, even with sparse feedback. This guarantee is extended to doubly robust estimators, enabling safety even under adversarial or misspecified user models and offering practitioners explicit control over permissible utility loss. The second part turns to single-action bandits, where various off-policy estimators are unified within a baseline-correction framework. A closed-form optimal baseline is proposed and shown to minimise both evaluation and policy-gradient variance, thereby improving off-policy learning reliability. The final part examines the trade-offs between efficiency and effectiveness in generative RL. A systematic study of PPO and REINFORCE motivates the Leave-One-Out PPO (LOOP) algorithm, which combines multiple diffusion trajectories with a REINFORCE-style baseline inside PPO's clipped objective. LOOP achieves PPO-level sample efficiency while producing generations that align more faithfully with textual attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。