arXiv:2510.12815cs.IR2025-10被引 4

用扩散模型提升推荐系统对长期用户偏好的预测能力

Energy-Guided Diffusion Sampling for Long-Term User Behavior Prediction in Reinforcement Learning-based Recommendation

  • 引入扩散过程建模用户偏好,增强离线强化学习的鲁棒性
  • 在6个真实数据集上验证,显著优化长期推荐效果
  • 适合研究离线强化学习推荐系统的学者和工程师

基于强化学习的推荐系统(RL4RS)因能适应动态用户偏好而受到关注。然而,这些系统在离线设置下面临数据效率低、依赖预收集轨迹的问题。尽管离线强化学习方法利用大规模数据缓解这些问题,但仍常受噪声数据影响,难以捕捉长期用户偏好,导致推荐策略不佳。为此,我们提出一种融合扩散过程与强化学习的新框架——DAC4Rec,通过扩散模型的去噪能力提升离线强化学习算法的鲁棒性,并采用基于Q值的策略优化机制以更好处理次优轨迹。此外,引入能量引导采样策略减少推荐生成中的随机性,确保结果更精准可靠。我们在六个真实世界离线数据集及在线模拟环境中进行了广泛实验,证明了该方法在优化长期用户偏好方面的有效性。进一步表明,所提出的扩散策略可无缝集成至其他主流RL4RS算法中,展现其通用性与广泛应用潜力。

原文摘要 · Abstract (English)

Reinforcement learning-based recommender systems (RL4RS) have gained attention for their ability to adapt to dynamic user preferences. However, these systems face challenges, particularly in offline settings, where data inefficiency and reliance on pre-collected trajectories limit their broader applicability. While offline reinforcement learning methods leverage extensive datasets to address these issues, they often struggle with noisy data and fail to capture long-term user preferences, resulting in suboptimal recommendation policies. To overcome these limitations, we propose Diffusion-enhanced Actor-Critic for Offline RL4RS (DAC4Rec), a novel framework that integrates diffusion processes with reinforcement learning to model complex user preferences more effectively. DAC4Rec leverages the denoising capabilities of diffusion models to enhance the robustness of offline RL algorithms and incorporates a Q-value-guided policy optimization strategy to better handle suboptimal trajectories. Additionally, we introduce an energy-based sampling strategy to reduce randomness during recommendation generation, ensuring more targeted and reliable outcomes. We validate the effectiveness of DAC4Rec through extensive experiments on six real-world offline datasets and in an online simulation environment, demonstrating its ability to optimize long-term user preferences. Furthermore, we show that the proposed diffusion policy can be seamlessly integrated into other commonly used RL algorithms in RL4RS, highlighting its versatility and wide applicability.

推荐系统扩散模型强化学习长期偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。