arXiv:2602.13852cs.AIstat.AP2026-02KDD

用AI自动选测试方案、解释胜败原因并生成创意建议。

Experimentation Accelerator: Interpretable Insights and Creative Recommendations for A/B Testing with Content-Aware ranking

  • 基于内容嵌入和历史数据,智能排序待测方案
  • 识别关键营销属性,定位高潜力未覆盖方向
  • 结合大模型生成可执行的创意优化建议

现代在线实验面临两大瓶颈:流量有限导致测试方案选择困难,事后分析依赖人工且常忽略内容特性。同时,组织未能充分利用历史实验结果与丰富的内容嵌入信息来指导优先级和创意迭代。本文提出统一框架,实现三方面能力:(i)智能推荐待测方案,(ii)解释胜出原因,(iii)发现新高潜力方案机会。通过引入处理项嵌入与历史结果,训练一个带固定效应的点击率(CTR)排名模型,平衡价值与内容多样性。为提升可解释性,将处理项投影至预定义语义营销属性空间,采用符号一致、稀疏约束的Lasso回归,获得每属性系数与符号贡献,用于可视化解释、关键驱动因素识别和自然语言洞察。进一步计算‘机会指数’,融合属性重要性与当前实验中该属性表达不足程度,标记缺失但高影响力的属性。最后,利用大语言模型(LLM)将排序后的机会转化为具体创意建议,并评估其学习与转化潜力,加速高效实验循环。该框架已集成至Adobe实际产品Experimentation Accelerator中,面向客户提供AI驱动的实验洞察与建议。通过真实业务实验验证,生成流程具备高质量表现。

原文摘要 · Abstract (English)

Modern online experimentation faces two bottlenecks: scarce traffic forces tough choices on which variants to test, and post-hoc insight extraction is manual, inconsistent, and often content-agnostic. Meanwhile, organizations underuse historical A/B results and rich content embeddings that could guide prioritization and creative iteration. We present a unified framework to (i) prioritize which variants to test, (ii) explain why winners win, and (iii) surface targeted opportunities for new, higher-potential variants. Leveraging treatment embeddings and historical outcomes, we train a CTR ranking model with fixed effects for contextual shifts that scores candidates while balancing value and content diversity. For better interpretability and understanding, we project treatments onto curated semantic marketing attributes and re-express the ranker in this space via a sign-consistent, sparse constrained Lasso, yielding per-attribute coefficients and signed contributions for visual explanations, top-k drivers, and natural-language insights. We then compute an opportunity index combining attribute importance (from the ranker) with under-expression in the current experiment to flag missing, high-impact attributes. Finally, LLMs translate ranked opportunities into concrete creative suggestions and estimate both learning and conversion potential, enabling faster, more informative, and more efficient test cycles. These components have been built into a real Adobe product, called \textit{Experimentation Accelerator}, to provide AI-based insights and opportunities to scale experimentation for customers. We provide an evaluation of the performance of the proposed framework on some real-world experiments by Adobe business customers that validate the high quality of the generation pipeline.

A/B测试可解释性创意生成AI实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。