arXiv:2505.12334cs.AI2025-05IJCAI被引 8

用批评者指导对话,让聊天机器人更主动了解用户喜好。

Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance

  • 引入批评者模型评估并引导对话的用户导向主动性
  • 构建新数据集ISCO-800,模拟多样化用户背景
  • 采用渐进式学习策略,适配不同沟通能力的用户

开放域对话系统旨在生成自然且吸引人的对话,在社交机器人和个人助手等实际应用中具有重要意义。大语言模型(LLMs)虽显著提升了上下文理解与对话流畅性,但现有系统仍难以主动捕捉用户聊天偏好并引导话题向用户中心倾斜。这种缺乏用户导向主动性的缺陷易使用户感到被忽视,降低满意度与继续对话意愿。为此,本文提出用户导向主动聊天机器人(UPC),首先借鉴LLM-as-a-judge思想构建批评者模型以评估主动性能;针对高质量训练数据稀缺问题,利用该批评者指导聊天机器人与用户代理之间的对话,生成具备更高用户导向主动性的语料库;为增强用户背景多样性,引入ISCO-800数据集构建用户代理;同时,针对用户沟通难度差异,提出迭代式课程学习方法,从易沟通用户逐步过渡到更难应对的用户,从而渐进提升聊天机器人表现。实验表明,该训练方法适用于多种LLM,在开放域对话中有效提升了用户导向主动性和对话吸引力。

原文摘要 · Abstract (English)

Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding and conversational fluency. However, existing LLM-based dialogue systems often fall short in proactively understanding the user's chatting preferences and guiding conversations toward user-centered topics. This lack of user-oriented proactivity can lead users to feel unappreciated, reducing their satisfaction and willingness to continue the conversation in human-computer interactions. To address this issue, we propose a User-oriented Proactive Chatbot (UPC) to enhance the user-oriented proactivity. Specifically, we first construct a critic to evaluate this proactivity inspired by the LLM-as-a-judge strategy. Given the scarcity of high-quality training data, we then employ the critic to guide dialogues between the chatbot and user agents, generating a corpus with enhanced user-oriented proactivity. To ensure the diversity of the user backgrounds, we introduce the ISCO-800, a diverse user background dataset for constructing user agents. Moreover, considering the communication difficulty varies among users, we propose an iterative curriculum learning method that trains the chatbot from easy-to-communicate users to more challenging ones, thereby gradually enhancing its performance. Experiments demonstrate that our proposed training method is applicable to different LLMs, improving user-oriented proactivity and attractiveness in open-domain dialogues.

对话系统大模型主动交互用户建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。