arXiv:2606.14046cs.IR2026-06

提出新方法缓解推荐去噪中的流行度偏差,保护长尾内容

When Recommendation Denoising Meets Popularity Bias: Understanding and Mitigating Their Interaction

论文配图:When Recommendation Denoising Meets Popularity Bias: Understanding and Mitigating Their Interaction
图 1 · 摘自论文原文
  • 根据物品流行度调节去噪强度,高曝光项强去噪,长尾项保守处理
  • 实验表明该方法在3个数据集上提升精度与多样性平衡
  • 特别适合对长尾内容敏感的MF类推荐模型

隐式反馈是推荐系统的主要数据来源,但行为日志常受误点击、曝光偏差和界面效应影响而含噪声。去噪方法通过降低或过滤疑似噪声交互来增强鲁棒性,通常依赖小损失启发式。本文从流行度偏差视角重新审视该启发式:冷门物品因观测稀疏,即使反映真实偏好也易产生较大损失,导致小损失重加权策略错误抑制此类干净但难拟合的信号,加剧头尾监督不平衡。我们形式化该交互关系,推导出条件再分配结果:当冷门正样本损失分布右移于热门项时,小损失重加权会使有效头尾信号比高于标准ERM。基于此,提出流行度感知去噪(PAD),一种轻量级插件框架,按物品流行度动态调节去噪强度——高曝光项强去噪,冷门项保留更多清洁信号。在三个数据集和三种骨干模型上的实验显示,PAD普遍优于代表性去噪基线,并在精度-多样性权衡上表现更优,尤其适用于MF类推荐器。

原文摘要 · Abstract (English)

Implicit feedback is the dominant data source for recommender systems, but behavioral logs are often contaminated by false-positive interactions caused by mis-clicks, biased exposure, and interface effects. Denoising recommendation methods improve robustness by down-weighting or filtering interactions suspected to be noisy, often relying on the small-loss heuristic. We revisit this heuristic through the lens of popularity bias. Tail-item positives can be harder to fit because they are sparsely observed, and thus may receive larger losses even when they reflect genuine user preference. Under such popularity-dependent loss patterns, monotone loss-based reweighting can suppress clean-but-hard tail signals and increase the head-tail imbalance in effective supervision. We formalize this interaction through the effective head-tail signal ratio induced by denoising weights and derive a conditional reallocation result: when the loss distribution of tail positives is right-shifted relative to that of head positives, small-loss reweighting increases the effective head-tail signal ratio compared with ERM. Motivated by this analysis, we propose Popularity-Aware Denoising (PAD), a lightweight plug-in framework that modulates denoising strength by item popularity. PAD applies stronger denoising to highly exposed items while being more conservative on tail items, preserving more clean-but-hard long-tail signals. Experiments on three datasets and three backbones show that PAD generally improves over representative denoising baselines and provides favorable accuracy-diversity tradeoffs, especially on MF-style recommenders.

推荐系统去噪流行度偏差长尾

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。