arXiv:2605.29405cs.LG2026-05

用信息导向策略提升离线到在线强化学习的探索效率。

Information-Directed Offline-to-Online Reinforcement Learning

  • 基于条件互信息设计信息导向采样,平衡即时损失与信息增益。
  • 在已知动态线性奖励模型中实现近似最优的贝叶斯后悔上界。
  • 适合离线数据有偏差但仍有价值的场景,如黑箱优化与贝叶斯优化。

从离线数据进行决策通常以固定离线数据预热策略或评分模型,再通过有限在线交互进行优化。离线数据虽减少不确定性,但无法消除探索需求;它仅改变剩余需探索的内容。我们通过条件互信息 $I(χ;τ_{1:T}ig| ext{D}_N)$ 形式化这种残余不确定性,其中 $χ$ 为学习目标,$τ_{1:T}$ 为在线轨迹,$\text{D}_N$ 为离线数据集。由此自然导出信息导向采样(IDS),其参数 $η ≥ 0$ 调控动作选择中即时后悔与信息收益的权衡。我们证明了 IDS 的通用贝叶斯后悔界:任何参考托马斯采样策略满足的信息比界,均可被 IDS 继承。在已知动力学的贝叶斯线性奖励模型中,条件互信息具对数行列式形式,且原版 IDS($η=0$)满足 $\widetilde O\left(Hd\min\left\{\sqrt T,\,T\sqrt{C^\dagger_{β,\mathrm{IDS}_0}(N,T)/N}\right\}\right)$,其中覆盖系数 $C^\dagger$ 与 IDS 本身诱导的访问分布相关。我们还发现一种预热阶段,其中原版 IDS 会选择具有信息量的探测动作,而托马斯采样永不如此,实现常数因子的贝叶斯后悔分离。受控扳机实验与 D4RL 离线到在线强化学习实验验证该机制:当离线数据提供信息但留下有偏或低概率的残余不确定性时,靶向在线动作最有效,此情形常见于离线强化学习、离线黑箱优化与贝叶斯优化。

原文摘要 · Abstract (English)

Decision-making from offline datasets typically warm-starts a policy or score model from fixed offline data and then refines it with limited online interaction. Offline data reduces uncertainty, but it does not remove the need for exploration; it changes what remains to be explored. We formalise this residual uncertainty by the conditional mutual information $I(χ;τ_{1:T}\mid\mathcal{D}_N)$ between a learning target $χ$ and the online trajectories after conditioning on the offline dataset. This view leads naturally to information-directed sampling (IDS), a family parameterised by $η\ge 0$ that selects actions by trading off instantaneous regret against information gain. We prove a generic offline-to-online Bayesian regret bound for IDS through a ratio certificate: any information-ratio bound satisfied by a reference Thompson-sampling policy over the same randomised policy class is inherited by IDS. In a known-dynamics Bayesian linear-reward model, the conditional mutual information has a log-determinant form, and vanilla IDS ($η=0$) satisfies $\widetilde O\!\left(Hd\min\left\{\sqrt T,\,T\sqrt{C^\dagger_{β,\mathrm{IDS}_0}(N,T)/N}\right\}\right),$ where the coverage coefficient is tied to the visitation distribution induced by vanilla IDS itself. We also identify a warm-start regime with a dominated but informative probe in which vanilla IDS selects the probe while Thompson sampling never does, giving a constant-factor Bayesian regret separation. Controlled bandit experiments and D4RL offline-to-online RL experiments validate this mechanism: IDS is most beneficial when offline data is informative but leaves biased or low-probability residual uncertainty that targeted online actions can resolve, a regime shared by offline RL, offline black-box optimization, and Bayesian optimization.

强化学习信息理论离线学习贝叶斯优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。