arXiv:2606.23933eess.SYcs.LG2026-06

通过修正历史数据漂移,让旧数据更有效用于动态环境下的决策

Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits

  • 用显式漂移模型把过去奖励迁移到当前,按可靠性加权重用
  • 在多类非平稳场景中性能超越传统丢弃旧数据的方法,尤其在周期性结构下提升显著
  • 适合处理有规律变化的动态系统,如金融投资、推荐系统等

我们研究奖励模型随时间漂移的非平稳线性上下文老虎机问题,传统方法因历史数据产生系统偏差而失效。提出流动修正汤普森采样(fcTS),一种贝叶斯方法:通过显式漂移模型将过去奖励迁移到当前,并对每个迁移观测赋予反映迁移可靠性的置信权重。该方法统一适用于三种情形:(i) 线性参数漂移,通过在线斜率估计与奖励修正;(ii) 周期性变化,基于相位感知跨周期重用;(iii) 重复的模式切换,结合变点检测与特定模式后验记忆。在线性高斯模型下后验更新保持闭式解,可通过截断且增量更新的充分统计量高效实现。在五个受控案例和一个包含多重重叠非平稳性的半合成投资组合选择基准上,fcTS均优于标准遗忘基线(折扣、滑动窗口、周期重启),在具有重复时间结构的场景中表现最佳。结果表明,当非平稳性具有结构时,修正并加权历史观测比统一丢弃更高效。

原文摘要 · Abstract (English)

We study non-stationary linear contextual bandits where the reward model drifts over time, rendering classical contextual bandit algorithms brittle because historical data becomes systematically biased. We propose Flow-Corrected Thompson Sampling (fcTS), a Bayesian method that reuses experience by transporting past rewards to the present using an explicit drift model and incorporating each transported observation with a confidence weight that reflects transport reliability. This yields a unified template that specializes in (i) linear parameter drift via online slope estimation and reward correction, (ii) periodic variation via phase-aware reuse across cycles, and (iii) recurring regime switches via changepoint detection and regime-specific posterior memory. The resulting posterior updates remain closed-form under a linear Gaussian model and can be implemented efficiently with truncated, incrementally updated sufficient statistics. Across five controlled case studies and a semi-synthetic portfolio-selection benchmark with multiple overlapping non-stationarities, fcTS outperforms standard forgetting-based baselines (discounting, sliding windows, and periodic restarts), with the largest gains in settings exhibiting recurring temporal structure. These results demonstrate that when non-stationarity is structured, correcting and reweighting historical observations can be substantially more sample-efficient than uniformly discarding them.

强化学习上下文老虎机非平稳性贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。