arXiv:2506.18744cs.LGstat.ML2025-06KDD被引 4

用快实验+慢实验结合,快速优化长期效果。

Experimenting, Fast and Slow: Bayesian Optimization of Long-term Outcomes with Online Experiments

  • 融合短期快实验与长期慢实验,用贝叶斯方法迭代优化。
  • 在短时间内完成对复杂系统策略的长期效果优化。
  • 适合需要快速决策但关注长期影响的在线系统优化场景。

互联网系统中的在线实验(即A/B测试)广泛用于系统调优,例如优化推荐系统的排序策略和自适应流媒体控制器。决策者通常希望优化系统变更的长期处理效果,但由于处理效果随时间非平稳,短期测量可能具有误导性,因此需要长时间运行实验。传统顺序实验策略需多轮迭代,耗时过长。本文提出一种新方法:将短期快实验(如仅运行数小时或数天的偏差实验)和/或离线代理(如脱策略评估)与长期慢实验相结合,实现大动作空间内高效、贝叶斯序贯优化,在极短时间内完成长期效果优化。

原文摘要 · Abstract (English)

Online experiments in internet systems, also known as A/B tests, are used for a wide range of system tuning problems, such as optimizing recommender system ranking policies and learning adaptive streaming controllers. Decision-makers generally wish to optimize for long-term treatment effects of the system changes, which often requires running experiments for a long time as short-term measurements can be misleading due to non-stationarity in treatment effects over time. The sequential experimentation strategies--which typically involve several iterations--can be prohibitively long in such cases. We describe a novel approach that combines fast experiments (e.g., biased experiments run only for a few hours or days) and/or offline proxies (e.g., off-policy evaluation) with long-running, slow experiments to perform sequential, Bayesian optimization over large action spaces in a short amount of time.

A/B测试贝叶斯优化长期效果在线实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。