arXiv:2506.14439cs.LG2025-06ICLR被引 1

用辅助奖励提升部分可观测目标下的离线强化学习效果

A General Framework for Off-Policy Learning with Partially-Observed Reward

  • 结合稀疏目标奖励与密集辅助奖励,设计新优化框架
  • 在真实数据和合成数据上均显著优于现有方法
  • 适合推荐系统、医疗决策等存在延迟或缺失反馈的场景

上下文关联的离线强化学习旨在仅使用历史数据学习最大化目标奖励的策略。然而,当奖励部分可观测时(如推荐系统的显式评分、电商转化信号延迟、医疗中的删失问题),传统方法性能急剧下降。一种应对策略是利用更密集观测的辅助奖励(如停留时长、点击、医学指标)。但仅依赖辅助奖励可能导致策略偏差。本文提出通用框架 HyPeR,通过融合部分观测的目标奖励与密集辅助奖励,实现高效离线学习。我们进一步探讨同时优化目标与辅助奖励的情形,发现此举反而有助于目标奖励优化。理论分析与实验证明,该方法在多种场景下均优于现有方法。

原文摘要 · Abstract (English)

Off-policy learning (OPL) in contextual bandits aims to learn a decision-making policy that maximizes the target rewards by using only historical interaction data collected under previously developed policies. Unfortunately, when rewards are only partially observed, the effectiveness of OPL degrades severely. Well-known examples of such partial rewards include explicit ratings in content recommendations, conversion signals on e-commerce platforms that are partial due to delay, and the issue of censoring in medical problems. One possible solution to deal with such partial rewards is to use secondary rewards, such as dwelling time, clicks, and medical indicators, which are more densely observed. However, relying solely on such secondary rewards can also lead to poor policy learning since they may not align with the target reward. Thus, this work studies a new and general problem of OPL where the goal is to learn a policy that maximizes the expected target reward by leveraging densely observed secondary rewards as supplemental data. We then propose a new method called Hybrid Policy Optimization for Partially-Observed Reward (HyPeR), which effectively uses the secondary rewards in addition to the partially-observed target reward to achieve effective OPL despite the challenging scenario. We also discuss a case where we aim to optimize not only the expected target reward but also the expected secondary rewards to some extent; counter-intuitively, we will show that leveraging the two objectives is in fact advantageous also for the optimization of only the target reward. Along with statistical analysis of our proposed methods, empirical evaluations on both synthetic and real-world data show that HyPeR outperforms existing methods in various scenarios.

离线学习奖励稀疏辅助信号推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。