用多目标强化学习打破推荐系统的信息茧房
Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation
- 将推荐建模为语义多目标决策过程,分别优化点击、多样性与公平性
- 在MovieLens小数据集上实现多样性提升,同时仅轻微影响用户点击率
- 适合关注算法公平性与信息生态健康的推荐系统研究者
推荐系统常因单一追求短期用户参与度而引发信息茧房与语义同质化。传统单目标模型难以平衡平台留存与信息多样性、提供方公平性等社会价值。为此,本文提出一种基于语义嵌入与帕累托深度Q网络(Pareto-DQN)的多目标强化学习框架,将点击率、多样性与公平性作为独立不可合并的奖励信号处理,避免静态加权带来的偏差。在MovieLens small数据集上的实证表明,基于超体积的动作选择能有效打断导致语义退化的反馈循环。通过保持高状态轨迹方差,该框架成功逼近帕累托前沿,在辅助目标上获得显著提升,对点击率影响微乎其微。本工作为构建内在对齐、负责任的推荐系统提供了可行路径。
原文摘要 · Abstract (English)
Recommender systems often induce filter bubbles and semantic homogenization by monolithically optimizing for immediate user engagement. Standard single-objective models, including traditional Deep Q-Networks, are ill-equipped to navigate the trade-offs between platform retention and critical societal values like information diversity and provider fairness. To address these limitations, we introduce a multi-objective reinforcement learning framework that formalizes recommendation as a semantic multi-objective Markov decision process. By integrating high-fidelity semantic embeddings with a Pareto-DQN agent, our architecture treats engagement, diversity, and fairness as distinct, non-aggregable reward signals, avoiding the pitfalls of static reward scalarization. Empirical evaluations on the MovieLens small dataset shows that our hypervolume based action selection disrupts the feedback loops responsible for semantic collapse. By sustaining high state-trajectory variance, the Pareto-DQN effectively maps the Pareto frontier, achieving gains in auxiliary societal objectives with only marginal impacts on engagement. This work provides a path toward intrinsically aligned, responsible recommender systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。