用扩散模型净化用户偏好状态,让推荐系统更公平
Fairness Begins with State: Purifying Latent Preferences for Hierarchical Reinforcement Learning in Interactive Recommendation
- 用扩散模型从噪声交互中还原真实用户偏好
- 在两个模拟器上实现推荐效果与公平性的更好平衡
- 适合关注推荐系统公平性与长期用户体验的研究者
交互式推荐系统(IRS)正越来越多地采用强化学习(RL)来捕捉用户-系统交互的序列特性。然而,现有公平性方法常忽略一个根本问题:观察到的用户状态并非真实偏好的忠实表示。隐式反馈受热门内容噪声和曝光偏差污染,导致状态失真,误导RL代理。我们认为,准确率与公平性之间的矛盾不仅是奖励设计问题,更是状态估计失败。本文提出DSRM-HRL框架,将公平性推荐重构为潜在状态净化问题,再进行解耦的层级决策。基于扩散模型的去噪状态表示模块(DSRM)从高熵、噪声化的交互历史中恢复低熵的潜在偏好流形。在此净化状态基础上,采用层级强化学习(HRL)代理解耦冲突目标:高层策略调控长期公平性轨迹,低层策略在动态约束下优化短期参与度。在高保真模拟器KuaiRec和KuaiRand上的大量实验表明,DSRM-HRL有效打破‘富者愈富’的反馈循环,在推荐效用与曝光公平性之间达成更优的帕累托前沿。
原文摘要 · Abstract (English)
Interactive recommender systems (IRS) are increasingly optimized with Reinforcement Learning (RL) to capture the sequential nature of user-system dynamics. However, existing fairness-aware methods often suffer from a fundamental oversight: they assume the observed user state is a faithful representation of true preferences. In reality, implicit feedback is contaminated by popularity-driven noise and exposure bias, creating a distorted state that misleads the RL agent. We argue that the persistent conflict between accuracy and fairness is not merely a reward-shaping issue, but a state estimation failure. In this work, we propose \textbf{DSRM-HRL}, a framework that reformulates fairness-aware recommendation as a latent state purification problem followed by decoupled hierarchical decision-making. We introduce a Denoising State Representation Module (DSRM) based on diffusion models to recover the low-entropy latent preference manifold from high-entropy, noisy interaction histories. Built upon this purified state, a Hierarchical Reinforcement Learning (HRL) agent is employed to decouple conflicting objectives: a high-level policy regulates long-term fairness trajectories, while a low-level policy optimizes short-term engagement under these dynamic constraints. Extensive experiments on high-fidelity simulators (KuaiRec, KuaiRand) demonstrate that DSRM-HRL effectively breaks the "rich-get-richer" feedback loop, achieving a superior Pareto frontier between recommendation utility and exposure equity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。