arXiv:2603.03843stat.MLcs.LG2026-03

通过挖掘奖励模型的不变性,提升非平稳线性强化学习的在线性能。

Invariance-Based Dynamic Regret Minimization

  • 将奖励模型分解为平稳与非平稳分量,利用历史数据中的不变性信息
  • 理论与实验表明,该方法可降低问题维度,显著减少快速变化环境下的累积损失
  • 适合有充足历史数据且环境动态变化的推荐系统、自适应控制等场景

我们研究随机非平稳线性带宽问题,其中上下文与回报之间的线性参数随时间变化。现有算法通过逐渐丢弃或降权历史数据来局部化策略,有效缩短了学习的时间窗口。但在许多场景中,历史数据仍可能携带关于回报模型的部分信息。本文提出利用这些信息以适应变化,假设回报模型可分解为平稳和非平稳成分。基于此,我们引入ISD-linUCB算法,利用历史数据学习回报模型中的不变性,并加以利用以提升在线表现。理论与实证均表明,借助不变性可降低问题维度,在历史数据充足时,快速变化环境中能实现显著的累积遗憾下降。

原文摘要 · Abstract (English)

We consider stochastic non-stationary linear bandits where the linear parameter connecting contexts to the reward changes over time. Existing algorithms in this setting localize the policy by gradually discarding or down-weighting past data, effectively shrinking the time horizon over which learning can occur. However, in many settings historical data may still carry partial information about the reward model. We propose to leverage such data while adapting to changes, by assuming the reward model decomposes into stationary and non-stationary components. Based on this assumption, we introduce ISD-linUCB, an algorithm that uses past data to learn invariances in the reward model and subsequently exploits them to improve online performance. We show both theoretically and empirically that leveraging invariance reduces the problem dimensionality, yielding significant regret improvements in fast-changing environments when sufficient historical data is available.

在线学习带宽算法非平稳性不变性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。