arXiv:2509.21501cs.HCcs.CL2025-09被引 15

用AI数字分身模拟用户,评估购物助手表现。

LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?

  • 用LLM Agent构建用户数字分身,复现真实交互
  • 数字分身行为模式与真人高度一致,反馈相似
  • 为快速评估智能助手提供可扩展的新方法

Agentic AI正迅速发展,能通过自然语言执行任务,如代码助手Copilot或购物助手Amazon Rufus。传统人工评估难以跟上其迭代速度。本文招募40名真实用户使用Amazon Rufus购物,收集其人物画像、交互轨迹和用户体验反馈,并创建对应的数字分身重复该任务。成对比较显示,尽管数字分身探索选项更广,但其行为模式与真人高度一致,且给出的设计反馈相似。这是首个量化评估LLM Agent在多轮交互中模仿人类能力的研究,证实其在可扩展系统评估中的潜力。

原文摘要 · Abstract (English)

Agentic AI is emerging, capable of executing tasks through natural language, such as Copilot for coding or Amazon Rufus for shopping. Evaluating these systems is challenging, as their rapid evolution outpaces traditional human evaluation. Researchers have proposed LLM Agents to simulate participants as digital twins, but it remains unclear to what extent a digital twin can represent a specific customer in multi-turn interaction with an agentic AI system. In this paper, we recruited 40 human participants to shop with Amazon Rufus, collected their personas, interaction traces, and UX feedback, and then created digital twins to repeat the task. Pairwise comparison of human and digital-twin traces shows that while agents often explored more diverse choices, their action patterns aligned with humans and yielded similar design feedback. This study is the first to quantify how closely LLM agents can mirror human multi-turn interaction with an agentic AI system, highlighting their potential for scalable evaluation.

Agentic AI用户模拟评估方法LLM Agent

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。