arXiv:2502.00747cs.CLcs.AI2025-02AAAI被引 1

用统一网络同时优化对话系统所有模块输出,提升任务完成率。

Universal Post-Processing Networks for Joint Optimization of Modules in Task-Oriented Dialogue Systems

  • 设计通用后处理网络,统一处理各模块输出
  • 在MultiWOZ上任务完成率显著优于传统方法
  • 适合需要端到端优化的对话系统研究者

后处理网络(PPNs)通过强化学习优化任务导向对话系统中任意模块的输出,以提升整体任务完成能力。然而,以往方法仅能处理系统中部分模块,限制了性能提升。本文提出通用后处理网络(UniPPN),基于语言模型实现对任意模块输出的序列化修改,并采用模块级马尔可夫决策过程的强化学习算法,实现对各模块的细粒度价值与优势估计,从而稳定联合优化过程。在MultiWOZ数据集上的仿真与人工评估实验表明,UniPPN在任务完成率方面显著优于传统PPN方法。

原文摘要 · Abstract (English)

Post-processing networks (PPNs) are components that modify the outputs of arbitrary modules in task-oriented dialogue systems and are optimized using reinforcement learning (RL) to improve the overall task completion capability of the system. However, previous PPN-based approaches have been limited to handling only a subset of modules within a system, which poses a significant limitation in improving the system performance. In this study, we propose a joint optimization method for post-processing the outputs of all modules using universal post-processing networks (UniPPNs), which are language-model-based networks that can modify the outputs of arbitrary modules in a system as a sequence-transformation task. Moreover, our RL algorithm, which employs a module-level Markov decision process, enables fine-grained value and advantage estimation for each module, thereby stabilizing joint learning for post-processing the outputs of all modules. Through both simulation-based and human evaluation experiments using the MultiWOZ dataset, we demonstrated that UniPPN outperforms conventional PPNs in the task completion capability of task-oriented dialogue systems.

对话系统强化学习后处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。