用生成式模拟用户自动对话,提升推荐系统理解复杂偏好能力
Search-Based Interaction For Conversation Recommendation via Generative Reward Model Based Simulated User
- 设计生成式反馈机制,模拟用户进行多轮交互
- 在公开数据集上提升推荐准确率,交互效率更高
- 适合需要高效个性化推荐的场景,如电商、内容平台
对话式推荐系统(CRS)通过多轮交互捕捉用户偏好,但用户偏好复杂多变,传统方法难以精准建模。为减少用户频繁参与带来的体验下降,本文提出基于生成式奖励模型的模拟用户(GRSU),通过自动反馈帮助系统更深入理解用户需求。该模拟用户支持两类反馈:粗粒度评分与细粒度属性批评,并统一为指令格式,通过合成数据指令微调实现一体化建模。结合束搜索策略优化交互过程,同时提出高效候选排序方法。在多个公开数据集上的实验表明,该方法在推荐效果、效率和可迁移性方面均表现优异。
原文摘要 · Abstract (English)
Conversational recommendation systems (CRSs) use multi-turn interaction to capture user preferences and provide personalized recommendations. A fundamental challenge in CRSs lies in effectively understanding user preferences from conversations. User preferences can be multifaceted and complex, posing significant challenges for accurate recommendations even with access to abundant external knowledge. While interaction with users can clarify their true preferences, frequent user involvement can lead to a degraded user experience. To address this problem, we propose a generative reward model based simulated user, named GRSU, for automatic interaction with CRSs. The simulated user provides feedback to the items recommended by CRSs, enabling them to better capture intricate user preferences through multi-turn interaction. Inspired by generative reward models, we design two types of feedback actions for the simulated user: i.e., generative item scoring, which offers coarse-grained feedback, and attribute-based item critique, which provides fine-grained feedback. To ensure seamless integration, these feedback actions are unified into an instruction-based format, allowing the development of a unified simulated user via instruction tuning on synthesized data. With this simulated user, automatic multi-turn interaction with CRSs can be effectively conducted. Furthermore, to strike a balance between effectiveness and efficiency, we draw inspiration from the paradigm of reward-guided search in complex reasoning tasks and employ beam search for the interaction process. On top of this, we propose an efficient candidate ranking method to improve the recommendation results derived from interaction. Extensive experiments on public datasets demonstrate the effectiveness, efficiency, and transferability of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。