arXiv:2602.01853cs.LGstat.ME2026-02被引 1

用Transformer+强化学习优化时间序列A/B测试,提升政策评估效果。

Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning

  • 用Transformer捕捉历史全貌,动态调整实验分配策略。
  • 在合成数据、网约车模拟器和真实数据上均显著降低误差。
  • 适合需要精准评估时序政策的平台型系统研究者。

A/B测试已成为科技公司评估政策的标准方法,但在时间序列实验中,即政策随时间顺序分配的情况下,仍面临挑战。现有设计存在两大局限:(i) 未充分利用完整历史信息进行处理分配;(ii) 为优化设计而依赖强假设来近似目标函数(如治疗效应估计的均方误差)。我们首先建立不可能定理,证明忽略完整历史会导致次优设计,因时间序列实验中存在动态依赖。为同时解决上述问题,我们提出一种基于Transformer强化学习(RL)的方法,利用Transformer对全历史条件化分配,并通过强化学习直接优化均方误差,无需强假设。在合成数据、公开的调度模拟器及真实网约车数据集上的实证评估表明,本方法始终优于现有设计。

原文摘要 · Abstract (English)

A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where policies are sequentially assigned over time, remains challenging. Existing designs suffer from two limitations: (i) they do not fully leverage the entire history for treatment allocation; (ii) they rely on strong assumptions to approximate the objective function (e.g., the mean squared error of the estimated treatment effect) for optimizing the design. We first establish an impossibility theorem showing that failure to condition on the full history leads to suboptimal designs, due to the dynamic dependencies in time series experiments. To address both limitations simultaneously, we next propose a transformer reinforcement learning (RL) approach which leverages transformers to condition allocation on the entire history and employs RL to directly optimize the MSE without relying on restrictive assumptions. Empirical evaluations on synthetic data, a publicly available dispatch simulator, and a real-world ridesharing dataset demonstrate that our proposal consistently outperforms existing designs.

A/B测试时间序列强化学习政策评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。