针对线性上下文猜拳问题,提出方向感知的离线-在线学习方法。
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits

- 引入方向性偏差证书,按特征方向区分离线数据的有用与误导程度
- 新算法在偏差方向对齐时可实现比标准方法更低的后悔值
- 适用于离线数据与在线环境存在差异但部分信息仍可用的场景
许多带宽系统部署时依赖离线历史数据(如早期策略的日志)。利用这些数据可减少在线探索成本,但若离线与在线环境不同,数据可能带来偏差。对于线性(上下文)带宽问题,这种偏差具有方向性:某些特征方向上数据仍有效,另一些则可能误导。以往工作通常通过已知的欧几里得参数界控制偏差,我们证明这过于粗糙——即使已知离线参数,单个未知方向的偏差仍会导致依赖维度的后悔增长。为此,我们提出方向性偏差证书 $(M_{\mathrm{bias}},ρ)$,通过 $M_{\mathrm{bias}}$-诱导范数衡量离线到在线的差距,并为不同方向分配不同偏差预算。基于此证书,我们设计 extit{Ellipsoidal-MINUCB} 算法,引入一个安全利用历史数据的离线融合分支。当证书已知时,算法最坏情况下达到标准 SupLinUCB 的后悔率,且在离线覆盖与低偏差方向一致时表现更优。当证书未知时,我们从离线和累积在线数据中自适应估计它,并建立相应后悔界。数值实验支持理论结果,在对齐场景下展现出显著收益。
原文摘要 · Abstract (English)
Many bandit systems are deployed with offline historical data, such as past logs from earlier policies. Using these data can reduce early online exploration when they remain informative for the online problem. When the offline and online environments differ, such data can be biased for the online problem. For linear (contextual) bandits, this bias is directional: offline data may be informative in some feature directions and misleading in others. However, prior work typically controls this gap through a known Euclidean bound on the model parameters, which we prove is too coarse: even with the offline parameter known, bias in a single unknown direction can force dimension-dependent regret. To address this challenge, we introduce a directional bias certificate $(M_{\mathrm{bias}},ρ)$ that measures the offline-to-online gap through an $M_{\mathrm{bias}}$-induced norm and assigns different bias budgets to different directions. Building on this certificate, we propose \emph{Ellipsoidal-MINUCB}, which augments the online learning with an offline-pooled branch that safely exploits historical data. When the certificate is known, we show that the algorithm matches the standard SupLinUCB rate in the worst case and improves when offline coverage aligns with low-bias directions. When the certificate is unknown, we estimate it adaptively from offline and accumulated online data and establish a corresponding regret guarantee. Numerical experiments support the theory and show gains in aligned regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。