arXiv:2509.20381cs.CLcs.AI2025-09中稿 · Recsys'25

让大模型学会对话推荐,训练推理双管齐下

USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model

  • 用模拟用户构建强化学习数据集,教大模型对话推荐策略
  • 推理阶段引入自增强机制,进一步挖掘推荐潜力
  • 在多个数据集上超越现有最佳方法,适合对话系统研究者

大语言模型(LLM)已广泛应用于对话推荐系统(CRS)。现有基于LLM的方法多聚焦于推理阶段的总结与分析能力,忽视了模型训练环节。为此,本文提出一个融合训练与推理的框架——用户模拟器驱动框架(USB-Rec),从模型层面提升LLM在对话推荐中的表现。首先,设计基于LLM的偏好优化(PO)数据集构建策略,用于强化学习训练,使LLM掌握对话推荐的方法与策略;其次,在推理阶段提出自增强策略(SES),进一步挖掘强化学习所获得的推荐潜力。在多个数据集上的大量实验表明,该方法持续优于此前最先进方法。

原文摘要 · Abstract (English)

Recently, Large Language Models (LLMs) have been widely employed in Conversational Recommender Systems (CRSs). Unlike traditional language model approaches that focus on training, all existing LLMs-based approaches are mainly centered around how to leverage the summarization and analysis capabilities of LLMs while ignoring the issue of training. Therefore, in this work, we propose an integrated training-inference framework, User-Simulator-Based framework (USB-Rec), for improving the performance of LLMs in conversational recommendation at the model level. Firstly, we design a LLM-based Preference Optimization (PO) dataset construction strategy for RL training, which helps the LLMs understand the strategies and methods in conversational recommendation. Secondly, we propose a Self-Enhancement Strategy (SES) at the inference stage to further exploit the conversational recommendation potential obtained from RL training. Extensive experiments on various datasets demonstrate that our method consistently outperforms previous state-of-the-art methods.

对话推荐大模型强化学习用户模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。