arXiv:2603.22335cs.IRcs.AI2026-03

用因果机制提升大模型推荐的泛化能力,避免环境干扰导致的偏差。

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

  • 引入因果不变性学习,通过反向通道调整消除环境混淆因子影响。
  • 在四种分布偏移设置下平均提升17.17%的推荐性能。
  • 适合关注推荐系统鲁棒性与跨环境泛化的研究者。

直接偏好优化(DPO)通过最小化偏好对齐损失,引导大语言模型生成与用户历史行为分布一致的推荐。然而,我们的系统性实证研究与理论分析表明,DPO在对齐过程中容易放大由环境混淆因子引起的虚假相关性,显著削弱基于大模型的生成式推荐方法在分布外(OOD)场景下的泛化能力。为此,我们提出CausalDPO,一种扩展的DPO方法,其引入因果不变性学习机制,在偏好对齐阶段采用反向通道调整策略以消除环境混淆因子的干扰,利用软聚类方法显式建模潜在环境分布,并通过不变性约束增强不同环境间的推荐一致性。理论分析表明,CausalDPO能有效捕捉用户在多环境中的稳定偏好结构,从而提升基于大模型推荐模型的OOD泛化性能。我们在四种代表性分布偏移设置下进行了大量实验,验证了CausalDPO的有效性,在四个评估指标上平均性能提升17.17%。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior distributions by minimizing preference alignment loss. However, our systematic empirical research and theoretical analysis reveal that DPO tends to amplify spurious correlations caused by environmental confounders during the alignment process, significantly undermining the generalization capability of LLM-based generative recommendation methods in out of distribution (OOD) scenarios. To mitigate this issue, we propose CausalDPO, an extension of DPO that incorporates a causal invariance learning mechanism. This method introduces a backdoor adjustment strategy during the preference alignment phase to eliminate interference from environmental confounders, explicitly models the latent environmental distribution using a soft clustering approach, and enhances robust consistency across diverse environments through invariance constraints. Theoretical analysis demonstrates that CausalDPO can effectively capture users stable preference structures across multiple environments, thereby improving the OOD generalization performance of LLM-based recommendation models. We conduct extensive experiments under four representative distribution shift settings to validate the effectiveness of CausalDPO, achieving an average performance improvement of 17.17% across four evaluation metrics.

推荐系统因果学习大模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。