arXiv:2504.09723cs.HCcs.CL2025-04被引 26

用AI代理模拟用户行为,实现快速可扩展的网页A/B测试。

AgentA/B: Automated and Scalable Web A/BTesting with Interactive LLM Agents

  • 用大模型代理自动模拟真实用户在网页上的多步操作。
  • 1000个代理在亚马逊网站测试,结果与真人行为高度相似。
  • 适合需要快速验证设计的互联网产品团队使用。

A/B测试是现代Web应用评估界面与用户体验设计的常用方法。然而传统A/B测试受限于对大规模真实用户流量的依赖以及漫长的等待周期。通过对六位行业专家的访谈,我们识别出当前工作流中的关键瓶颈。为此,我们提出AgentA/B,一个基于大语言模型的自主代理系统,可自动模拟用户与真实网页的交互行为。AgentA/B支持以多样化角色规模化部署LLM代理,每名代理能动态导航网页并执行搜索、点击、筛选、购买等多步交互。在一项受控实验中,我们在Amazon.com上使用1,000个代理进行跨被试组的A/B测试,并将代理行为与真实人类购物行为在规模上进行对比。结果表明,AgentA/B能有效模拟类人行为模式。

原文摘要 · Abstract (English)

A/B testing experiment is a widely adopted method for evaluating UI/UX design decisions in modern web applications. Yet, traditional A/B testing remains constrained by its dependence on the large-scale and live traffic of human participants, and the long time of waiting for the testing result. Through formative interviews with six experienced industry practitioners, we identified critical bottlenecks in current A/B testing workflows. In response, we present AgentA/B, a novel system that leverages Large Language Model-based autonomous agents (LLM Agents) to automatically simulate user interaction behaviors with real webpages. AgentA/B enables scalable deployment of LLM agents with diverse personas, each capable of navigating the dynamic webpage and interactively executing multi-step interactions like search, clicking, filtering, and purchasing. In a demonstrative controlled experiment, we employ AgentA/B to simulate a between-subject A/B testing with 1,000 LLM agents Amazon.com, and compare agent behaviors with real human shopping behaviors at a scale. Our findings suggest AgentA/B can emulate human-like behavior patterns.

A/B测试大模型代理自动化用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。