arXiv:2508.07243cs.LGcs.AI2025-08被引 2

用扩散模型生成无偏负样本,提升推荐系统在分布外场景的泛化能力。

Causal Negative Sampling via Diffusion Model for Out-of-Distribution Recommendation

  • 通过条件扩散过程在隐空间合成负样本,避开候选池中的环境偏差
  • 在四种分布偏移场景下,平均性能提升13.96%,显著优于现有方法
  • 适合需要强鲁棒性的推荐系统,尤其适用于数据分布变化大的场景

启发式负样本采样通过从预定义候选池中选择不同难度的负样本,引导模型学习更准确的决策边界,从而提升推荐性能。然而,我们的实证与理论分析表明,候选池中存在的未观测环境混杂因素(如曝光或流行度偏差)可能导致启发式方法引入虚假硬负样本(FHNS)。这些误导性样本会促使模型学习由混杂因素引发的虚假相关性,最终损害其在分布偏移下的泛化能力。为此,我们提出一种名为因果负采样通过扩散(CNSDiff)的新方法。该方法通过条件扩散过程在隐空间合成负样本,避免了预定义候选池带来的偏差,从而降低生成虚假硬负样本的可能性。此外,其引入因果正则化项,显式抑制环境混杂因素对负采样过程的影响,生成更具鲁棒性的负样本,促进分布外(OOD)泛化。在四种典型分布偏移场景下的综合实验表明,与最先进基线相比,CNSDiff在所有评估指标上平均提升13.96%,验证了其在分布外推荐任务中的有效性和鲁棒性。

原文摘要 · Abstract (English)

Heuristic negative sampling enhances recommendation performance by selecting negative samples of varying hardness levels from predefined candidate pools to guide the model toward learning more accurate decision boundaries. However, our empirical and theoretical analyses reveal that unobserved environmental confounders (e.g., exposure or popularity biases) in candidate pools may cause heuristic sampling methods to introduce false hard negatives (FHNS). These misleading samples can encourage the model to learn spurious correlations induced by such confounders, ultimately compromising its generalization ability under distribution shifts. To address this issue, we propose a novel method named Causal Negative Sampling via Diffusion (CNSDiff). By synthesizing negative samples in the latent space via a conditional diffusion process, CNSDiff avoids the bias introduced by predefined candidate pools and thus reduces the likelihood of generating FHNS. Moreover, it incorporates a causal regularization term to explicitly mitigate the influence of environmental confounders during the negative sampling process, leading to robust negatives that promote out-of-distribution (OOD) generalization. Comprehensive experiments under four representative distribution shift scenarios demonstrate that CNSDiff achieves an average improvement of 13.96% across all evaluation metrics compared to state-of-the-art baselines, verifying its effectiveness and robustness in OOD recommendation tasks.

推荐系统因果学习扩散模型分布外泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。