arXiv:2410.12519cs.IR2024-10被引 16

让大模型推荐更符合用户真实偏好,同时减少偏见和幻觉。

RosePO: Aligning LLM-based Recommenders with Human Values

  • 通过优化选择与拒绝样本,提升推荐的个性化和帮助性。
  • 在三个真实数据集上,推荐效果更好,且降低流行度偏差和语义幻觉。
  • 适合关注推荐系统公平性与可信度的研究者与开发者。

近年来,人们越来越关注将大语言模型(LLM)用于推荐系统,通常通过监督微调(SFT)将预训练模型适配到推荐场景。然而,预训练和SFT阶段均未显式建模用户对不同物品偏好的比较关系。为构建“有益且无害”的基于LLM的推荐系统,我们提出通用框架RosePO——带平滑个性化偏好优化的推荐。该框架在后训练阶段更好地对齐个性化人类价值观。具体而言,除自然对齐于SFT数据的输入与优选回复外,我们设计了针对提升帮助性的拒绝采样策略,以及两种缓解偏见以促进无害性的策略。为应对自动构建偏好数据中不确定标签的问题,我们在优化目标中引入由偏好预言机预测的个性化平滑因子。在三个真实世界数据集上的评估表明,该方法不仅提升了推荐性能,还有效缓解了语义幻觉和流行度偏差。

原文摘要 · Abstract (English)

Recently, there has been a growing interest in leveraging Large Language Models (LLMs) for recommendation systems, which usually adapt a pre-trained LLM to the recommendation scenario through supervised fine-tuning (SFT). However, both the pre-training and SFT stages fail to explicitly model the comparative relationships of a user's preferences on different items. To construct a "helpful and harmless" LLM-based recommender, we propose a general framework -- Recommendation with smoothing personalized Preference Optimization (RosePO), which better aligns with customized human values during the post-training stage. Specifically, in addition to the input and chosen response that naturally align with SFT data, we design a rejected sampling strategy tailored for enhancing helpfulness, along with two strategies aimed at mitigating biases to promote harmlessness. To ensure robustness against uncertain labels present in automatically constructed preference data, we introduce a personalized smoothing factor predicted by a preference oracle into the optimization objective. Evaluation on three real-world datasets demonstrates the effectiveness of our method, showcasing not only improved recommendation performance but also mitigation of semantic hallucination and popularity bias.

推荐系统大模型偏好优化无害性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。