arXiv:2505.01592cs.CLcs.AI2025-05被引 1

提出AURA框架,用用户交互过程诊断智能体满意度

AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents

  • 构建交互式规划智能体的行为阶段评估框架
  • 发现用户满意度受中间行为与最终结果共同影响
  • 适合研究人机交互、智能体评估的学者与开发者

大型语言模型在指令遵循和上下文理解方面能力增强,推动了具备复杂内部流程(如上下文理解、工具管理、响应生成)的智能体广泛应用。然而,现有基准多以任务完成率作为整体效能代理指标,我们提出假设:仅提升任务完成率并不等于最大化用户满意度,因为用户关注整个交互过程而非仅结果。为此,我们提出AURA(Agent-User inteRaction Assessment)框架,将交互式任务规划智能体的行为阶段概念化,并通过一组原子级大模型评估标准,全面评估智能体在决策流程中的表现。分析表明,不同智能体在各行为阶段表现各异,用户满意度由最终结果及中间行为共同决定。我们还指出未来方向,包括多智能体系统应用及用户模拟器在任务规划中的局限性。

原文摘要 · Abstract (English)

The growing capabilities of large language models (LLMs) in instruction-following and context-understanding lead to the era of agents with numerous applications. Among these, task planning agents have become especially prominent in realistic scenarios involving complex internal pipelines, such as context understanding, tool management, and response generation. However, existing benchmarks predominantly evaluate agent performance based on task completion as a proxy for overall effectiveness. We hypothesize that merely improving task completion is misaligned with maximizing user satisfaction, as users interact with the entire agentic process and not only the end result. To address this gap, we propose AURA, an Agent-User inteRaction Assessment framework that conceptualizes the behavioral stages of interactive task planning agents. AURA offers a comprehensive assessment of agent through a set of atomic LLM evaluation criteria, allowing researchers and practitioners to diagnose specific strengths and weaknesses within the agent's decision-making pipeline. Our analyses show that agents excel in different behavioral stages, with user satisfaction shaped by both outcomes and intermediate behaviors. We also highlight future directions, including systems that leverage multiple agents and the limitations of user simulators in task planning.

智能体评估人机交互用户满意度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。