arXiv:2508.05289cs.LG2025-08被引 21

用隐式反馈训练对话推荐系统,让模型更懂用户真实偏好。

RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders

  • 通过强化学习融合隐式反馈信号优化对话推荐
  • 在REDIAL和OpenDialKG上提升推荐准确率与用户满意度
  • 适合需要动态适配用户偏好的对话推荐场景

基于大语言模型的对话推荐系统(CRS)需持续对齐用户偏好以提供满意且上下文相关的推荐。传统监督微调无法捕捉隐式反馈信号,如停留时间、情感极性或参与模式。本文提出一种基于人类反馈强化学习(RLHF)的微调方案,旨在多轮推荐场景中最大化隐式用户反馈(IUF)。我们构建了一个在弱标注互动信息上训练的奖励模型 $R_ϕ$,并通过近端策略优化(PPO)方法优化基础语言模型 $M_θ$,以提升用户中心效用。该架构建模对话状态转移 $s_t \to a_t \to s_{t+1}$,其中动作 $a_t$ 仅根据历史对话生成物品建议。在合成数据及真实数据集(如 REDIAL、OpenDialKG)上的评估表明,相比基线模型,本方法在 top-$k$ 推荐准确率、连贯性与用户满意度方面均有显著提升。结果表明,隐式信号对齐可在实现可扩展、用户自适应的 CRS 设计中发挥高效作用。

原文摘要 · Abstract (English)

Conversational recommender systems (CRS) based on Large Language Models (LLMs) need to constantly be aligned to the user preferences to provide satisfying and context-relevant item recommendations. The traditional supervised fine-tuning cannot capture the implicit feedback signal, e.g., dwell time, sentiment polarity, or engagement patterns. In this paper, we share a fine-tuning solution using human feedback reinforcement learning (RLHF) to maximize implied user feedback (IUF) in a multi-turn recommendation context. We specify a reward model $R_ϕ$ learnt on weakly-labelled engagement information and maximize user-centric utility by optimizing the foundational LLM M_θ through a proximal policy optimization (PPO) approach. The architecture models conversational state transitions $s_t \to a_t \to s_{t +1}$, where the action $a_t$ is associated with LLM-generated item suggestions only on condition of conversation history in the past. The evaluation across synthetic and real-world datasets (e.g.REDIAL, OpenDialKG) demonstrates that our RLHF-fine-tuned models can perform better in terms of top-$k$ recommendation accuracy, coherence, and user satisfaction compared to (arrow-zero-cmwrquca-teja-falset ensuite 2Round group-deca States penalty give up This paper shows that implicit signal alignment can be efficient in achieving scalable and user-adaptive design of CRS.

对话推荐隐式反馈RLHF大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。