用模拟器生成对话数据,提升招聘对话代理的主动引导能力
SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection
- 构建高保真用户模拟器,生成大规模多轮对话数据
- 基于意图链评估框架,筛选高质量训练数据,提升效果
- 在真实招聘场景中验证,适合工业级对话系统开发
任务导向的主动对话代理在招聘中至关重要,尤其在引导对话达成特定业务目标(如获取社交媒体联系方式以转入私密频道)方面。尽管监督微调和强化学习已被证明有效,但其性能受限于高质量、目标导向的领域特定数据稀缺。为此,我们提出SimRPD,一种三阶段训练框架:首先,开发高保真用户模拟器,通过多轮在线对话生成大规模对话数据;其次,引入基于意图链(Chain-of-Intention, CoI)的多维度评估框架,综合全局与实例级指标,全面评估模拟器并高效筛选高质量数据;最后,基于筛选后的数据训练招聘主动对话代理。在真实招聘场景中的实验表明,SimRPD优于现有基于模拟器的数据选择策略,凸显其在工业部署中的实用价值,并具备向其他业务导向对话场景扩展的潜力。
原文摘要 · Abstract (English)
Task-oriented proactive dialogue agents play a pivotal role in recruitment, particularly for steering conversations towards specific business outcomes, such as acquiring social-media contacts for private-channel conversion. Although supervised fine-tuning and reinforcement learning have proven effective for training such agents, their performance is heavily constrained by the scarcity of high-quality, goal-oriented domain-specific training data. To address this challenge, we propose SimRPD, a three-stage framework for training recruitment proactive dialogue agents. First, we develop a high-fidelity user simulator to synthesize large-scale conversational data through multi-turn online dialogue. Then we introduce a multi-dimensional evaluation framework based on Chain-of-Intention (CoI) to comprehensively assess the simulator and effectively select high-quality data, incorporating both global-level and instance-level metrics. Finally, we train the recruitment proactive dialogue agent on the selected dataset. Experiments in a real-world recruitment scenario demonstrate that SimRPD outperforms existing simulator-based data selection strategies, highlighting its practical value for industrial deployment and its potential applicability to other business-oriented dialogue scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。