arXiv:2605.22240cs.AI2026-05

让对话代理主动发现用户隐性需求,提升销售类对话说服力

Unlocking Proactivity in Task-Oriented Dialogue

论文配图:Unlocking Proactivity in Task-Oriented Dialogue
图 1 · 摘自论文原文
  • 用分层人格模型模拟用户隐性关切,生成可追踪说服进度的对话
  • 通过双视角优化,使模型在有限轮次内主动引导用户接受
  • 适合需要高说服力的智能客服、销售机器人场景

主动型任务导向对话(如外呼销售)要求对话代理在有限轮次内主动探查用户隐性关切并引导其接受。然而,后训练大模型本质上保守,基于奖励塑形的强化学习(如GRPO)因仅重加权已有被动策略的采样而效果受限。本文表明,以用户隐性关切为条件可激活主动能力,且该能力无法被采样过程削弱,确立其作为关键训练信号的重要性。为此,我们构建了认知用户模拟器,将每位用户建模为包含可观测外部特征与隐藏内部关切的分层人格。该模拟器生成真实且多样化的交互,并输出每轮的状态动态以追踪说服进展。进一步提出模拟器诱导的非对称视图策略优化:(1) 非对称在线策略自蒸馏,将同一策略的特权视角(含关切信息)中的关切感知行为迁移到仅基于对话的部署视角;(2) 状态转移策略精炼……

原文摘要 · Abstract (English)

Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the conversation toward acceptance within a bounded number of turns. Yet post-trained LLMs are inherently conservative, and reward-shaping RL (e.g., GRPO) struggles since it only re-weights what an already passive policy samples. We show that conditioning on the user's latent concerns unlocks proactive capability that no amount of sampling can undermine, establishing these concerns as a pivotal training-time signal. To operationalize this finding, we build the \textbf{Cognitive User Simulator}, which models each user as a stratified persona comprising observable external traits and hidden internal concerns. The simulator produces faithful and diverse interactions, while emitting per-turn state dynamics that track persuasion progress. We then introduce \textbf{Simulator-Induced Asymmetric-View Policy Optimization}, which converts the modeled concerns and the simulation state transition into complementary training objectives: (1) \emph{Asymmetric On-Policy Self-Distillation} that transfers concern-aware behavior from a privileged view of the same policy into its deployable, conversation-only view; and (2) \emph{State-Transition Policy Refinement} ...

对话系统主动对话强化学习用户建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。