对比真人与智能体在搜索系统中的行为差异,发现结果相似但路径不同。
Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search Systems
- 从任务结果、查询方式、界面跳转三方面追踪比较人类与智能体行为
- 智能体成功率与真人相当,查询内容也相近,但导航路径更机械
- 适合关注智能体行为真实性与部署可信度的研究者和工程师
基于大模型的GUI智能体正被广泛用于生产环境,以自动化流程并模拟用户进行评估与优化。然而现有评估多关注任务完成率,缺乏对智能体是否具备类人行为的深入验证。本文提出一种基于轨迹的评估框架,从任务结果与投入、查询生成方式、界面状态间跳转行为三方面,对比真实用户与智能体表现。我们在一个真实的音频流搜索应用中开展对照实验,39名参与者与当前最先进的GUI智能体共同完成10个需多跳检索的任务。结果显示,智能体任务成功率与人类相当,查询内容高度一致,但在导航策略上存在系统性差异:人类表现出以内容为中心的探索性行为,而智能体则更偏向以搜索为中心且分支较少的路径。这表明任务结果与查询一致性不等于行为一致,提醒在将智能体作为用户代理部署时,必须引入细粒度行为诊断。
原文摘要 · Abstract (English)
LLM-driven GUI agents are increasingly used in production systems to automate workflows and simulate users for evaluation and optimization. Yet most GUI-agent evaluations emphasize task success and provide limited evidence on whether agents interact in human-like ways. We present a trace-level evaluation framework that compares human and agent behavior across (i) task outcome and effort, (ii) query formulation, and (iii) navigation across interface states. We instantiate the framework in a controlled study in a production audio-streaming search application, where 39 participants and a state-of-the-art GUI agent perform ten multi-hop search tasks. The agent achieves task success comparable to participants and generates broadly aligned queries, but follows systematically different navigation strategies: participants exhibit content-centric, exploratory behavior, while the agent is more search-centric and low-branching. These results show that outcome and query alignment do not imply behavioral alignment, motivating trace-level diagnostics when deploying GUI agents as proxies for users in production search systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。