用精英个体注入提升对话策略的探索效率
An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite Individuals
- 融合进化算法与强化学习,兼顾全局搜索与局部优化
- 在四个数据集上显著提升对话性能,收敛更快
- 适合需要高效训练对话系统的研究者和工程师
深度强化学习(DRL)广泛应用于任务导向型对话系统以优化对话策略,但其在高维状态与动作空间中难以平衡探索与利用,常陷入局部最优或收敛缓慢。进化算法(EAs)通过保持种群多样性,能有效探索神经网络解空间。受此启发,本文创新性地结合EA的全局搜索能力与DRL的局部优化优势,实现探索与利用的平衡。然而,自然语言在对话任务中的固有灵活性使这种直接融合复杂化,导致进化时间过长。为此,我们进一步提出精英个体注入机制(EII),通过自适应地将表现最佳个体引入种群,提升进化搜索效率。在四个公开数据集上的实验表明,该方法显著改善了探索与利用的平衡,提升了性能;同时,EII机制有效缩短了探索时间,实现了EA与DRL在任务导向对话策略任务中的高效融合。
原文摘要 · Abstract (English)
Deep Reinforcement Learning (DRL) is widely used in task-oriented dialogue systems to optimize dialogue policy, but it struggles to balance exploration and exploitation due to the high dimensionality of state and action spaces. This challenge often results in local optima or poor convergence. Evolutionary Algorithms (EAs) have been proven to effectively explore the solution space of neural networks by maintaining population diversity. Inspired by this, we innovatively combine the global search capabilities of EA with the local optimization of DRL to achieve a balance between exploration and exploitation. Nevertheless, the inherent flexibility of natural language in dialogue tasks complicates this direct integration, leading to prolonged evolutionary times. Thus, we further propose an elite individual injection mechanism to enhance EA's search efficiency by adaptively introducing best-performing individuals into the population. Experiments across four datasets show that our approach significantly improves the balance between exploration and exploitation, boosting performance. Moreover, the effectiveness of the EII mechanism in reducing exploration time has been demonstrated, achieving an efficient integration of EA and DRL on task-oriented dialogue policy tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。