用大模型动态规划对话策略,零训练实现高精度目标对话
A General Highly Accurate Online Planning Method Integrating Large Language Models into Nested Rollout Policy Adaptation for Dialogue Tasks
- 用大模型同时模拟用户和系统行为,构建嵌套蒙特卡洛优化框架
- 在4个数据集上超越现有提示工程与预训练模型方法,性能接近ChatGPT
- 仅需0.6亿参数大模型,无需微调即可适应新对话场景
在目标导向对话任务中,核心挑战是在有限轮次内引导对话达成目标。现有方法或依赖复杂提示工程(效果依赖人工经验),或融合策略网络与预训练模型(难适应新场景且训练成本高)。本文提出一种新型对话策略规划方法NRPA-GD,完全避免特定模型训练,利用大语言模型(LLM)同时模拟用户与系统行为。该方法构建完整的对话轨迹评估机制,并采用嵌套蒙特卡洛仿真与策略自适应优化框架,在对话过程中动态调整策略。在四个典型目标导向对话数据集上的实验表明,NRPA-GD优于现有提示工程及预训练模型方法。令人印象深刻的是,仅使用0.6亿参数的LLM,其性能即超越ChatGPT与预训练策略模型。该方法进一步验证了将规划方法应用于大模型以解决实际任务的可行性与创新性。
原文摘要 · Abstract (English)
In goal-oriented dialogue tasks, the main challenge is to steer the interaction towards a given goal within a limited number of turns. Existing approaches either rely on elaborate prompt engineering, whose effectiveness is heavily dependent on human experience, or integrate policy networks and pre-trained policy models, which are usually difficult to adapt to new dialogue scenarios and costly to train. Therefore, in this paper, we present Nested Rollout Policy Adaptation for Goal-oriented Dialogue (NRPA-GD), a novel dialogue policy planning method that completely avoids specific model training by utilizing a Large Language Model (LLM) to simulate behaviors of user and system at the same time. Specifically, NRPA-GD constructs a complete evaluation mechanism for dialogue trajectories and employs an optimization framework of nested Monte Carlo simulation and policy self-adaptation to dynamically adjust policies during the dialogue process. The experimental results on four typical goal-oriented dialogue datasets show that NRPA-GD outperforms both existing prompt engineering and specifically pre-trained model-based methods. Impressively, NRPA-GD surpasses ChatGPT and pre-trained policy models with only a 0.6-billion-parameter LLM. The proposed approach further demonstrates the advantages and novelty of employing planning methods on LLMs to solve practical planning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。