arXiv:2411.05194cs.LGcs.AI2024-11被引 10

让对话智能体通过事后反思优化互动策略,提升说服与共情能力。

Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations

  • 基于事后回溯重构对话数据,用离线强化学习训练交互策略。
  • 在真实用户实验中显著优于现有顶尖对话模型,尤其在心理支持和募捐场景。
  • 无需专家标注,适合需要理解对方心理状态的高阶对话任务。

大型语言模型(LLMs)虽能生成自然流畅的对话文本,但当前方法多聚焦于单次准确应答,忽视了真实对话中的交互性——即说话者需通过言语影响对方、获取信息或改变其观点。为实现有效对话引导,现有方法依赖专家数据,但需理解对方认知过程,这超出了人类与普通训练模型的能力。本文提出关键洞察:虽模型难以预判或实时制定有效策略,但可在对话结束后,根据对方反馈反向优化自身表达。我们利用此机制重构并扩充低效对话数据,采用离线强化学习训练对话智能体,在心理健康支持与慈善募捐两个需理解人类心理状态的任务中表现优异。真实用户研究表明,该方法显著超越现有最先进模型。

原文摘要 · Abstract (English)

Recent progress on large language models (LLMs) has enabled dialogue agents to generate highly naturalistic and plausible text. However, current LLM language generation focuses on responding accurately to questions and requests with a single effective response. In reality, many real dialogues are interactive, meaning an agent's utterances will influence their conversational partner, elicit information, or change their opinion. Accounting for how an agent can effectively steer a conversation is a crucial ability in many dialogue tasks, from healthcare to preference elicitation. Existing methods for fine-tuning dialogue agents to accomplish such tasks would rely on curating some amount of expert data. However, doing so often requires understanding the underlying cognitive processes of the conversational partner, which is a skill neither humans nor LLMs trained on human data can reliably do. Our key insight is that while LLMs may not be adept at identifying effective strategies for steering conversations a priori, or in the middle of an ongoing conversation, they can do so post-hoc, or in hindsight, after seeing how their conversational partner responds. We use this fact to rewrite and augment existing suboptimal data, and train via offline reinforcement learning (RL) an agent that outperforms both prompting and learning from unaltered human demonstrations. We apply our approach to two domains that require understanding human mental state, intelligent interaction, and persuasion: mental health support, and soliciting charitable donations. Our results in a user study with real humans show that our approach greatly outperforms existing state-of-the-art dialogue agents.

对话系统强化学习心理支持交互策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。