用虚拟用户模拟A/B测试,快速评估设计优劣。
SimAB: Simulating A/B Tests with Persona-Conditioned AI Agents for Rapid Design Evaluation
- 用带角色设定的AI代理模拟真实用户行为
- 对47个历史实验预测准确率达67%,高置信度时达83%
- 适合低流量页面、隐私敏感场景的快速设计筛选
A/B测试是验证设计决策的标准方法,但依赖真实用户流量限制了迭代速度,使某些实验难以实施。我们提出SimAB,一种将A/B测试重构为快速、隐私保护的模拟系统,使用角色设定的AI代理。给定设计截图和转化目标,SimAB生成用户角色,部署代理表达偏好,聚合结果并生成推理理由。通过与实践者的形式化研究,我们识别出流量受限导致测试困难的场景,包括低流量页面、多变量比较、微调优化及隐私敏感情境。设计强调速度、早期反馈、可操作的推理和受众指定。我们在47个已知结果的历史A/B测试上评估SimAB,整体准确率达67%,高置信度情况下提升至83%。额外实验显示其对命名和位置偏差具有鲁棒性,并证明角色设定带来准确率提升。从业者反馈表明,SimAB支持更快的评估周期,能高效筛选传统A/B测试难以评估的设计。
原文摘要 · Abstract (English)
A/B testing is a standard method for validating design decisions, yet its reliance on real user traffic limits iteration speed and makes certain experiments impractical. We present SimAB, a system that reframes A/B testing as a fast, privacy-preserving simulation using persona-conditioned AI agents. Given design screenshots and a conversion goal, SimAB generates user personas, deploys them as agents that state their preference, aggregates results, and synthesizes rationales. Through a formative study with experimentation practitioners, we identified scenarios where traffic constraints hinder testing, including low-traffic pages, multi-variant comparisons, micro-optimizations, and privacy-sensitive contexts. Our design emphasizes speed, early feedback, actionable rationales, and audience specification. We evaluate SimAB against 47 historical A/B tests with known outcomes, achieving 67% overall accuracy, increasing to 83% for high-confidence cases. Additional experiments show robustness to naming and positional bias and demonstrate accuracy gains from personas. Practitioner feedback suggests that SimAB supports faster evaluation cycles and rapid screening of designs difficult to assess with traditional A/B tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。