arXiv:2606.13683cs.AIcs.CL2026-06

基于用户画像动态调整对话策略,提升目标导向对话成功率

UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems

论文配图:UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems
图 1 · 摘自论文原文
  • 根据用户画像实时调整对话策略,无需离线训练
  • 多任务成功率达100%,谈判任务成交比提升56.41%
  • 适合需要个性化交互的智能客服与谈判系统

为解决现有对话策略规划方法难以动态适应多样用户特征的问题,本文提出基于用户画像的嵌套回溯策略自适应(UP-NRPA)在线框架,结合大语言模型实现动态策略定制。与依赖模型训练和离线强化学习策略模型的传统方法不同,UP-NRPA通过实时用户反馈及从当前用户画像中提取的人格、偏好和目标,实现无需离线强化学习即可自适应用户特征的对话策略调整。在协作与非协作对话基准测试中表现显著,多个对话任务成功率达到100%;尤其在谈判任务中,销售对列表比率(SL)提升了56.41%。结果表明,UP-NRPA可在不依赖训练机制的情况下有效适配多样化用户需求,使对话系统具备更强的用户适应能力。

原文摘要 · Abstract (English)

To address the challenge that current dialogue policy planning methods struggle to dynamically adapt to diverse user characteristics, this paper proposes a User Portrait based Nested Rollout Policy Adaptation (UP-NRPA) online framework with Large Language Models. In contrast to conventional approaches dependent on model training and require offline reinforcement learning policy models for user groups, UP-NRPA enables dynamic customization of dialogue strategies through an adaptive mechanism. This is achieved by leveraging real-time user feedback alongside personality, preferences, and objectives mapped from the current user portrait, thereby adapting to user characteristics without offline reinforcement learning. In collaborative and non-collaborative dialogue benchmarks, UP-NRPA demonstrated considerable benefits, achieving an impressive 100% success rate in multiple dialogue tasks. Particularly in negotiation tasks, the sale-to-list ratio (SL) increased by 56.41%. This demonstrates that UP-NRPA can adapt to diverse user needs without requiring a training mechanism, enabling the dialogue system to adapt to user characteristics.

对话系统用户画像大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。