构建真实多轮交互数据,提升智能体训练效果。
PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems

- 基于四维人格空间与真实行为统计生成用户模拟
- 训练后综合得分提升4.1%,任务完成率增6.0%
- 适合需要真实多轮交互训练的智能体研究者
大语言模型越来越多地被用作代理工作流执行器,但现有训练数据和评估基准大多假设信息完整、单轮查询。对1.6万次真实会话的分析显示,75.9%的交互为多轮对话,暴露出用户与代理之间的互动方式与训练评估方式之间的显著差距。本文提出PersonaForge,一个用于合成真实多轮用户-代理交互的用户模拟框架。该框架结合四维人格空间、基于真实用户统计数据校准的SOUL驱动行为控制,以及基于真实种子查询的反向深度构造技术。利用PersonaForge,我们构建了一个包含6.3K条记录的训练数据集和名为PersonaForge-Bench的手动标注基准,涵盖超过20个专业领域,采用四维评分体系。在Qwen3.5-27B上的实验表明,PersonaForge训练使综合得分提升+4.1%,各维度均有增长,其中任务完成率提升+6.0%,响应质量提升+6.8%。进一步分析显示,训练后的代理使用更少轮次和工具调用,表明交互效率提高;消融实验证实了SOUL组件和自适应模拟的有效性。PersonaForge与PersonaForge-Bench共同建立了在真实多轮交互下训练与评估代理的基础。
原文摘要 · Abstract (English)
Large language models are increasingly used as agentic workflow executors, yet existing training data and benchmarks largely assume informationally complete, single-turn queries. Our analysis of 16K real-world sessions shows that 75.9% of interactions are multi-turn, revealing a substantial gap between how users interact with agents and how such systems are trained and evaluated. We introduce \textbf{PersonaForge}, a user simulation framework for synthesizing realistic multi-turn user--agent interactions. PersonaForge combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries. Using PersonaForge, we construct a 6.3K-record training dataset and \textbf{PersonaForge-Bench}, a manually annotated 138-task benchmark spanning over 20 professional domains with four-dimensional scoring. Experiments on Qwen3.5-27B show that PersonaForge training improves the composite score by +4.1%, with gains across all four dimensions and the largest improvements in Task Completion (+6.0%) and Response Quality (+6.8%). Further analyses show that PersonaForge-trained agents use fewer turns and tool calls, suggesting improved interaction efficiency, while ablations confirm the contribution of SOUL components and adaptive simulation. Together, PersonaForge and PersonaForge-Bench establish a foundation for training and evaluating agents under realistic multi-turn user interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。