让语言模型主动提问澄清意图,提升人机协作效率
SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
- 用强化学习奖励模型主动提问,增强对话主动性
- 任务完成率比基线高20.14%,且不增加对话轮次
- 适合需要精准交互的智能助手、客服系统等场景
高效的人机协作在现实应用中日益普遍。当前协作多为单向:用户下达指令或提问,模型直接回应,缺乏必要的澄清与确认。随着模型能力提升,需推动更主动的参与机制,使模型能动态提问以明确意图、解决歧义并适应变化。现有工作低估了语言模型的对话潜力,导致模型仅被优化为被动响应者。本文提出SpeakRL,一种基于强化学习的方法,通过奖励主动交互(如适时提问)来提升模型的对话能力。为此,我们构建了SpeakER合成数据集,涵盖任务导向对话中的多样化场景,任务通过交互式澄清问题解决。我们系统分析了对话主动性奖励设计,并提出一种平衡提问与执行的合理奖励公式。实证评估表明,该方法在任务完成率上相比基线模型绝对提升20.14%,且未增加对话轮次,甚至超越更大规模的私有模型,验证了以澄清为核心的交互模式的有效性。
原文摘要 · Abstract (English)
Effective human-agent collaboration is increasingly prevalent in real-world applications. Current trends in such collaborations are predominantly unidirectional, with users providing instructions or posing questions to agents, where agents respond directly without seeking necessary clarifications or confirmations. However, the evolving capabilities of these agents require more proactive engagement, where agents should dynamically participate in conversations to clarify user intents, resolve ambiguities, and adapt to changing circumstances. Existing prior work under-utilize the conversational capabilities of language models (LMs), thereby optimizing agents as better followers rather than effective speakers. In this work, we introduce SpeakRL, a reinforcement learning (RL) method that enhances agents' conversational capabilities by rewarding proactive interactions with users, such as asking right clarification questions when necessary. To support this, we curate SpeakER, a synthetic dataset that includes diverse scenarios from task-oriented dialogues, where tasks are resolved through interactive clarification questions. We present a systematic analysis of reward design for conversational proactivity and propose a principled reward formulation for teaching agents to balance asking with acting. Empirical evaluations demonstrate that our approach achieves a 20.14% absolute improvement in task completion over base models without increasing conversation turns even surpassing even much larger proprietary models, demonstrating the promise of clarification-centric user-agent interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。