用奖励引导课程采样提升意图识别泛化能力
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
- 基于奖励的课程采样在强化学习中聚焦难例训练
- 相比监督微调,泛化性能显著提升
- 适合需要应对未知意图的对话系统研发
意图识别是任务导向对话系统的核心组件,面临快速新增工具及其复杂关联关系带来的挑战。现有方法如零样本重构和基于大模型的动态识别,在遇到未见意图时性能下降,导致任务路由错误。为提升模型对未见任务的泛化能力,本文在分组相对策略优化(GRPO)训练中引入强化学习与基于奖励的课程采样(RCS)。实验表明,强化学习训练模型在泛化性能上显著优于监督微调基线;引入RCS有效增强强化学习效果,使模型聚焦于困难案例。此外,在强化学习中结合思维链(COT)过程,显著提升了复杂意图识别任务的泛化能力,凸显了思维过程在复杂场景中的重要性。本工作推进了意图识别的泛化能力,为可适应对话系统的部署提供实用洞见。
原文摘要 · Abstract (English)
Intent detection, a critical component in task-oriented dialogue (TOD) systems, faces significant challenges in adapting to the rapid influx of integrable tools with complex interrelationships. Existing approaches, such as zero-shot reformulations and LLM-based dynamic recognition, struggle with performance degradation when encountering unseen intents, leading to erroneous task routing. To enhance the model's generalization performance on unseen tasks, we employ Reinforcement Learning (RL) combined with a Reward-based Curriculum Sampling (RCS) during Group Relative Policy Optimization (GRPO) training in intent detection tasks. Experiments demonstrate that RL-trained models substantially outperform supervised fine-tuning (SFT) baselines in generalization. Besides, the introduction of the RCS, significantly bolsters the effectiveness of RL in intent detection by focusing the model on challenging cases during training. Moreover, incorporating Chain-of-Thought (COT) processes in RL notably improves generalization in complex intent detection tasks, underscoring the importance of thought in challenging scenarios. This work advances the generalization of intent detection tasks, offering practical insights for deploying adaptable dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。