arXiv:2609.02170cs.LG2026-09

用强化学习自动优化广告技能文档,提升推荐效果。

DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation

论文配图:DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation
图 1 · 摘自论文原文
  • 将文档编辑建模为序列动作,通过上下层代理协同优化。
  • 在短视频平台实测,多指标优于现有方法。
  • 擅长需要长期评估的广告策略迭代,适合算法工程师。

广告推荐需持续调优复杂系统参数,以平衡商业收益与用户体验。现有工作引入大语言模型(LLMs)与技能文档辅助此过程,但技能优化仍主要依赖提示驱动,缺乏对具体文档修改的奖励归因机制。为此,我们提出文档介导强化学习(DMRL),一个技能自演化框架,将技能文档优化建模为一系列结构化编辑动作。在DMRL中,上层智能体执行受控文档编辑,下层冻结的任务代理通过A/B测试评估其效果。为解决信用分配与长期结果问题,我们引入两个关键组件:(1) 双重相对策略优化(DRPO),一种后训练策略优化方法,实现鲁棒且风险感知的优势估计;(2) 长期奖励预测器(LRP),通过解耦表征学习与跨注意力迁移,建模用户群体异质性以预测长期结果。DMRL已在大规模短视频广告平台部署,大量实证评估表明,其在关键广告指标上均超越现有最先进基线。

原文摘要 · Abstract (English)

Advertising recommendation requires continuously tuning complex system parameters while balancing commercial returns and user experience. Recent work has introduced large language models (LLMs) with skill documents to assist this labor-intensive process, but skill optimization remains largely prompt-driven, lacking a principled mechanism to attribute rewards to specific document edits. To address this limitation, we propose Document-Mediated Reinforcement Learning (DMRL), a skill self-evolution framework that models skill document optimization as a sequence of structured editing actions. In DMRL, an upper-level agent performs controlled document edits, while a frozen lower-level task agent evaluates their effects through A/B testing. To address credit assignment and long-term outcomes, we introduce two key components: (1) Dual-Relative Policy Optimization (DRPO), a post-training policy optimization method for robust and risk-aware advantage estimation; and (2) Long-term Reward Predictor (LRP), which estimates long-term outcomes by modeling population heterogeneity with disentangled representation learning and cross-attention transfer. DMRL was deployed on a large-scale short-video ads platform and extensive empirical evaluation shows that DMRL outperforms state-of-the-art baselines across key advertising metrics

广告推荐强化学习大模型应用策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。