在动态变化的环境中,用历史数据预测未来政策效果并优化。
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
- 利用时间序列中的季节、周、节日等规律设计新型重要性加权方法
- 在非平稳环境下,对下月政策效果的估计误差显著低于现有方法
- 适合电商推荐等随时间变化场景,可仅用历史数据优化未来策略
我们研究了非平稳环境中未来离策略评估(F-OPE)与学习(F-OPL)的新问题,目标是利用过去策略收集的历史数据,估计并优化未来某时间点的策略价值。例如,在电商推荐中,需基于上月数据预测并优化下月策略表现。核心挑战在于未来环境的数据在历史中不可见。现有方法依赖平稳性假设或严苛的奖励建模假设,导致严重偏差。为此,我们提出首个利用时间结构的离策略未来价值估计器(OPFV),通过挖掘历史与未来共有的季节、周、节假日等时间模式,实现有效估计。理论分析揭示了其低偏差的条件。进一步,我们将其扩展为一种仅依赖历史数据的策略梯度学习方法,主动优化未来策略。实验表明,该方法在多种设置下均显著优于现有方法。
原文摘要 · Abstract (English)
We study the novel problem of future off-policy evaluation (F-OPE) and learning (F-OPL) for estimating and optimizing the future value of policies in non-stationary environments, where distributions vary over time. In e-commerce recommendations, for instance, our goal is often to estimate and optimize the policy value for the upcoming month using data collected by an old policy in the previous month. A critical challenge is that data related to the future environment is not observed in the historical data. Existing methods assume stationarity or depend on restrictive reward-modeling assumptions, leading to significant bias. To address these limitations, we propose a novel estimator named \textit{\textbf{O}ff-\textbf{P}olicy Estimator for the \textbf{F}uture \textbf{V}alue (\textbf{\textit{OPFV}})}, designed for accurately estimating policy values at any future time point. The key feature of OPFV is its ability to leverage the useful structure within time-series data. While future data might not be present in the historical log, we can leverage, for example, seasonal, weekly, or holiday effects that are consistent in both the historical and future data. Our estimator is the first to exploit these time-related structures via a new type of importance weighting, enabling effective F-OPE. Theoretical analysis identifies the conditions under which OPFV becomes low-bias. In addition, we extend our estimator to develop a new policy-gradient method to proactively learn a good future policy using only historical data. Empirical results show that our methods substantially outperform existing methods in estimating and optimizing the future policy value under non-stationarity for various experimental setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。