arXiv:2606.09802cs.LGcs.AI2026-06

解决用户偏好与环境漂移下的高效推荐问题,提升决策质量。

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

  • 设计Dri-MED算法应对偏好漂移和非平稳噪声,结合异方差回归。
  • 理论证明其后悔率与约束违反次数均优于现有方法。
  • 适合需实时调整策略的个性化推荐系统场景。

我们研究了一种线性上下文随机多臂赌博机变体,其中学习者需向具有个性化偏好向量的用户群体提供推荐,且上下文分布随时间漂移。在符合实践假设的前提下,该设置可简化为均值平稳但噪声异方差且非平稳的线性赌博机。进一步考虑学习者必须确保每一步决策的期望收益高于基准策略 $\boldsymbolπ_0$。为此,提出受启发于线性MED策略的Dri-MED算法,并针对非平稳异方差噪声进行了精细适配。理论分析表明,实例相关后悔率约为 $\tilde{\mathcal{O}}\left(\frac{κ}{\tilde{Δ}}d^2\log(T)\right)$,其中 $\tilde{Δ}$ 是受约束的次优间隙,$κ$ 为方差感知乘性项,通过异方差回归妥善处理。同时,证明了Dri-MED的期望约束违反次数为 $\tilde{\mathcal{O}}(d)$。数值实验显示,该算法显著优于忽略漂移和偏好结构的保守基线。

原文摘要 · Abstract (English)

We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time. Under practitioner-friendly assumptions, we reduce this setting to linear bandit with stationary mean but heteroskedastic and non-stationary noise. We further study the case when the learner must ensure the mean reward of each decision must exceed that of a baseline strategy $\boldsymbolπ_0$ at each decision step. We introduce Dri-MED, an algorithm inspired from the linear version of the MED strategy, and carefully adapted to handle the non-stationary heteroskedastic noise. We show that the instance-dependent regret scales as $\tilde{\mathcal O}\left(\fracκ{\tildeΔ}d^2(\log(T)\right)$, where $\tildeΔ$ is the constraint-aware sub-optimality gap subject to policy $π_0$, with variance-aware multiplicative term $κ$ that we carefully handle using heteroskedastic regression. We further show Dri-MED enjoys $\tilde{\mathcal{O}}(d)$ expected constraint violations. Our numerical results suggest that Dri-MED significantly outperforms conservative baselines that ignores the drift and preference structure.

强化学习在线决策个性化推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。