arXiv:2506.22112cs.IR2025-06中稿 · Companion Proceedi…被引 2

解决推荐系统离线强化学习中的奖励失衡问题,提升推荐多样性与模型准确性。

Reward Balancing Revisited: Enhancing Offline Reinforcement Learning for Recommender Systems

  • 利用模型不确定性重分配奖励,缓解预测偏差
  • 引入衰减惩罚机制,提升推荐多样性并优化全局状态覆盖
  • 适合追求个性化与多样性平衡的推荐系统研究者

离线强化学习已成为现实推荐系统中从历史数据中学习用户偏好、构建推荐策略的有效方法。然而,在离线强化学习中,奖励塑造面临严峻挑战:以往方法虽尝试通过先验信息处理不确定性或惩罚未充分探索的状态-动作对,但仍存在核心缺陷——难以同时平衡世界模型的内在偏差与策略推荐的多样性。为此,本文提出一种名为R3S(Reallocated Reward for Recommender Systems)的新型离线强化学习框架。该框架通过融合模型固有的不确定性来应对奖励预测的内在波动,增强决策多样性,以契合更交互式的推荐范式;并引入具有衰减特性的额外惩罚项,从局部与全局尺度抑制导致状态多样性下降的行为。实验结果表明,R3S不仅提升了世界模型的预测精度,还有效调和了用户群体中异质性偏好,显著改善推荐性能。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) has emerged as a prevalent and effective methodology for real-world recommender systems, enabling learning policies from historical data and capturing user preferences. In offline RL, reward shaping encounters significant challenges, with past efforts to incorporate prior strategies for uncertainty to improve world models or penalize underexplored state-action pairs. Despite these efforts, a critical gap remains: the simultaneous balancing of intrinsic biases in world models and the diversity of policy recommendations. To address this limitation, we present an innovative offline RL framework termed Reallocated Reward for Recommender Systems (R3S). By integrating inherent model uncertainty to tackle the intrinsic fluctuations in reward predictions, we boost diversity for decision-making to align with a more interactive paradigm, incorporating extra penalizers with decay that deter actions leading to diminished state variety at both local and global scales. The experimental results demonstrate that R3S improves the accuracy of world models and efficiently harmonizes the heterogeneous preferences of the users.

推荐系统离线RL奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。