arXiv:2608.15949cs.IRcs.AI2026-08

用信息熵衡量对话价值,让大模型更聪明地问问题。

Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

论文配图:Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
图 1 · 摘自论文原文
  • 以推荐熵减为奖励,优化对话策略
  • 在两个数据集上提升推荐质量与对话效率
  • 无需真实推荐答案,适合真实场景应用

大语言模型(LLM)已用于对话式推荐系统(CRS),表现出良好的推荐准确性和自然对话能力。然而,如何有效引导多轮对话以获取用户偏好仍具挑战。现有方法或依赖独立强化学习代理和模板化交互,或通过另一大模型评估互动性,却未衡量实际获得的信息量。本文提出新方法,通过推荐熵的减少来量化每轮交互的有效性,将熵减作为奖励信号,无需依赖真实推荐标签(此类标签在真实场景中常不可得),对大模型进行微调,从而生成更具策略性的对话。在INSPIRED和ReDial数据集上,采用监督微调(SFT)和直接偏好优化(DPO)的实验证明,该方法同时提升了推荐质量与对话效率。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue. However, guiding multi-turn interactions to elicit user preferences effectively remains challenging. Existing approaches either use separate reinforcement learning agents with templated interactions or optimize for interactivity judged by another LLM, without measuring how much useful information is actually gained. We propose a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations. We apply this entropy reduction as a reward---without relying on ground-truth recommendations, which are often unavailable in real-world scenarios---to fine-tune the LLM, enabling strategic interaction generation. Empirical results with supervised fine-tuning (SFT) and direct preference optimization (DPO) on the INSPIRED and ReDial datasets show that our method improves both recommendation quality and conversational efficiency.

对话推荐大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。