arXiv:2506.19652cs.CLcs.AI2025-06被引 1

用强化学习让大模型对话更智能,能自动调整策略适应不同用户。

Tailored Conversations beyond LLMs: A RL-Based Dialogue Manager

  • 分层强化学习建模对话阶段,元学习提升跨用户适应力
  • 在动机访谈任务中,奖励得分超越顶尖大模型基线
  • 适合需要个性化、目标导向对话的医疗或心理场景

本文提出一种融合大语言模型(LLMs)与基于强化学习的对话管理框架,用于具有特定目标的开放域对话。通过分层强化学习建模对话的结构化阶段,并采用元学习提升在多样化用户群体中的适应能力,该方法显著增强系统的适应性与效率,使其能在有限数据下学习、顺畅切换对话阶段,并个性化响应异构用户需求。我们在动机访谈任务中应用该框架,旨在促进行为改变,结果表明所提出的对话管理器在奖励指标上优于当前最先进的大模型基线,证明了对大模型进行目标引导以构建开放域对话系统具有潜力。

原文摘要 · Abstract (English)

In this work, we propose a novel framework that integrates large language models (LLMs) with an RL-based dialogue manager for open-ended dialogue with a specific goal. By leveraging hierarchical reinforcement learning to model the structured phases of dialogue and employ meta-learning to enhance adaptability across diverse user profiles, our approach enhances adaptability and efficiency, enabling the system to learn from limited data, transition fluidly between dialogue phases, and personalize responses to heterogeneous patient needs. We apply our framework to Motivational Interviews, aiming to foster behavior change, and demonstrate that the proposed dialogue manager outperforms a state-of-the-art LLM baseline in terms of reward, showing a potential benefit of conditioning LLMs to create open-ended dialogue systems with specific goals.

对话系统强化学习大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。