arXiv:2601.15290cs.HCcs.AI2026-01中稿 · NeurIPS被引 1

用多智能体模拟真实用户行为,提升对话系统测试的多样性与可信度。

Agentic Persona Control and Task State Tracking for Realistic User Simulation in Interactive Scenarios

  • 分三类智能体:用户调度、任务状态跟踪、对话属性生成。
  • 相比单模型基线,各项指标均显著提升,任务完成率更高。
  • 适合用于需要高仿真用户交互的对话系统评估场景。

大规模测试对话AI系统需涵盖多样领域的真实用户交互,以捕捉丰富的行为模式。本文提出一种新型多智能体框架,用于在交互场景中实现可解释、逼真的真人用户模拟,通过人格控制与任务状态追踪,模仿目标导向对话中的认知过程。系统包含三个专用智能体:(1) 用户代理负责整体交互调度,(2) 状态跟踪代理维护结构化任务状态,(3) 消息属性生成代理根据任务进展与分配人格控制对话特征。为验证方法有效性,我们在餐厅点餐场景下实现并评估该框架,覆盖高复杂度任务、多样化行为及对话模糊性。通过系统性消融实验,评估各组件对人格一致性、任务完成准确率、可解释性与真实感的贡献。实验表明,完整多智能体系统相较单大模型基线,在所有评估指标上均有显著提升。该框架构建了可分解、具认知合理性的智能体协同环境,适用于多种交互领域的人类用户模拟。

原文摘要 · Abstract (English)

Testing conversational AI systems at scale across diverse domains necessitates realistic and diverse user interactions capturing a wide array of behavioral patterns. We present a novel multi-agent framework for realistic, explainable human user simulation in interactive scenarios, using persona control and task state tracking to mirror human cognitive processes during goal-oriented conversations. Our system employs three specialized AI agents: (1) a User Agent to orchestrate the overall interaction, (2) a State Tracking Agent to maintain structured task state, and (3) a Message Attributes Generation Agent that controls conversational attributes based on task progress and assigned persona. To validate our approach, we implement and evaluate the framework for guest ordering at a restaurant with scenarios rich in task complexity, behavioral diversity, and conversational ambiguity. Through systematic ablations, we evaluate the contributory efficacy of each agentic component to overall simulation quality in terms of persona adherence, task completion accuracy, explainability, and realism. Our experiments demonstrate that the complete multi-agent system achieves superior simulation quality compared to single-LLM baselines, with significant gains across all evaluation metrics. This framework establishes a powerful environment for orchestrating agents to simulate human users with cognitive plausibility, decomposing the simulation into specialized sub-agents that reflect distinct aspects of human thought processes applicable across interactive domains.

用户模拟多智能体对话系统认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。