主动引导用户偏好,提升长尾商品曝光且不降低满意度。
Proactive Guiding Strategy for Item-side Fairness in Interactive Recommendation
- 用分层强化学习逐步引导用户偏好向长尾物品倾斜。
- 相比顶尖方法,交互奖励和用户停留时长显著提升。
- 适合关注推荐系统公平性与长期用户体验的开发者。
物品侧公平性对确保长尾物品在互动推荐系统中获得公平曝光至关重要。现有方法通过直接将长尾物品纳入推荐结果来提升其曝光,但这导致用户偏好与推荐内容错位,损害长期用户参与度并降低推荐效果。本文提出一种主动公平引导策略,旨在互动推荐过程中主动引导用户偏好向长尾物品靠拢,同时保持用户满意度。为此,我们提出HRL4PFG框架,利用分层强化学习实现对用户偏好的渐进式引导。该框架包含宏观层面:基于多步反馈生成公平引导目标;微观层面:实时结合目标与用户偏好变化精细调整推荐。大量实验表明,在互动推荐环境中,相比现有最先进方法,HRL4PFG在累积交互奖励和最大用户交互时长上均有更显著提升。
原文摘要 · Abstract (English)
Item-side fairness is crucial for ensuring the fair exposure of long-tail items in interactive recommender systems. Existing approaches promote the exposure of long-tail items by directly incorporating them into recommended results. This causes misalignment between user preferences and the recommended long-tail items, which hinders long-term user engagement and reduces the effectiveness of recommendations. We aim for a proactive fairness-guiding strategy, which actively guides user preferences toward long-tail items while preserving user satisfaction during the interactive recommendation process. To this end, we propose HRL4PFG, an interactive recommendation framework that leverages hierarchical reinforcement learning to guide user preferences toward long-tail items progressively. HRL4PFG operates through a macro-level process that generates fairness-guided targets based on multi-step feedback, and a micro-level process that fine-tunes recommendations in real time according to both these targets and evolving user preferences. Extensive experiments show that HRL4PFG improves cumulative interaction rewards and maximum user interaction length by a larger margin when compared with state-of-the-art methods in interactive recommendation environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。