用分阶段规划提升网络诱骗对话模拟的连贯性与目标达成率。
StagePilot: Stage-Level Planning for Long-Horizon Dialogue Simulation in Cybergrooming
- 将对话分为阶段,分步规划下一步动作并生成响应。
- 相比基线,43%相对提升到达最终阶段的概率,且正向回应超70%。
- 适合安全教育、儿童保护等需高可控性对话系统的研发者。
网络诱骗是威胁青少年的持续性风险,亟需主动干预。本文将对话发展建模为分阶段的结构化规划问题,提出StagePilot框架,通过分离阶段级规划与响应生成,使模型在约束转换下选择下一阶段,并基于该阶段生成回应,从而实现连贯且真实的对话推进。采用强化学习从离线数据中学习阶段策略,优化情绪契合度与目标一致性。实验证明,StagePilot生成的对话轨迹更结构化、更少停滞;尤其IQL+AWAC变体在更多情况下达成最终阶段,同时保持超过70%的正向或中性回应,相对改进达43%。
原文摘要 · Abstract (English)
Cybergrooming is an evolving threat to youth, requiring proactive educational interventions. We address this by modeling dialogue progression as a structured planning problem over stage-wise interactions. We propose StagePilot, a dialogue framework that separates stage-level planning from response generation, in which the model selects the next stage under constrained transitions and generates responses conditioned on it, enabling coherent and realistic progression. Reinforcement learning is used to learn stage-level policies from offline data, optimizing for both emotional alignment and goal-consistent progression. Our empirical experiments show that StagePilot generates more structured, coherent dialogue trajectories and reduces conversational stagnation compared to baselines; notably, the IQL+AWAC variant reaches the final stage more often while maintaining over 70% positive or neutral responses, yielding a 43% relative improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。