arXiv:2510.11997cs.CL2025-10Conference of the …被引 5

用商业知识构建真实用户模拟器,提升多轮对话智能体评估效果

SAGE: A Top-Down Bottom-Up Knowledge-Grounded User Simulator for Multi-turn AGent Evaluation

  • 结合业务逻辑与真实数据构建用户行为模型
  • 比现有方法多发现33%的智能体错误
  • 适合需要高仿真测试的对话系统研发团队

多轮交互智能体的评估面临人工评测成本高的挑战。现有用户模拟方法通常仅建模通用用户行为,忽视领域特性。本文提出SAGE框架,融合业务上下文知识进行用户模拟:上层基于理想客户画像等业务逻辑,生成符合真实客户角色的行为;下层引入产品目录、常见问题集和知识库等实际业务数据,使模拟交互更贴近目标市场的信息需求与预期。实验表明,该方法生成的交互更真实多样,同时能识别出比传统方法多33%的智能体缺陷,显著提升评估有效性,对漏洞发现与迭代优化具有重要价值。

原文摘要 · Abstract (English)

Evaluating multi-turn interactive agents is challenging due to the need for human assessment. Evaluation with simulated users has been introduced as an alternative, however existing approaches typically model generic users and overlook the domain-specific principles required to capture realistic behavior. We propose SAGE, a novel user Simulation framework for multi-turn AGent Evaluation that integrates knowledge from business contexts. SAGE incorporates top-down knowledge rooted in business logic, such as ideal customer profiles, grounding user behavior in realistic customer personas. We further integrate bottom-up knowledge taken from business agent infrastructure (e.g., product catalogs, FAQs, and knowledge bases), allowing the simulator to generate interactions that reflect users' information needs and expectations in a company's target market. Through empirical evaluation, we find that this approach produces interactions that are more realistic and diverse, while also identifying up to 33% more agent errors, highlighting its effectiveness as an evaluation tool to support bug-finding and iterative agent improvement.

对话系统用户模拟智能体评估知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。