arXiv:2502.07064cs.LGcs.AI2025-02NeurIPS被引 5

用生成模型补全缺失数据,让上下文强化学习更准更稳。

Contextual Thompson Sampling via Generation of Missing Data

  • 用生成模型随机填补缺失的未来或反事实结果
  • 理论证明可达到当前最优的后悔值上界
  • 适合需要可靠决策的推荐、医疗等场景

我们提出一种基于生成模型的上下文老虎机算法框架,其不确定性量化与决策能力取决于离线训练的生成模型质量。不同于传统将不确定性归因于不可观测隐变量,本方法认为不确定性源于潜在可观测但尚未出现的结果(包括未来和反事实结果)。若所有结果都可观测,便可直接用‘神谕’策略在完整数据集上拟合。受此启发,每次决策时,算法利用生成模型对缺失结果进行概率性填补,基于补全后的数据拟合策略,并据此选择动作。我们形式化证明该算法是汤普森采样的生成式表述,并建立了当前最优的后悔界。值得注意的是,该后悔界仅依赖生成模型的离线预测误差,且适用于任意“神谕”策略拟合方法。

原文摘要 · Abstract (English)

We introduce a framework for Thompson sampling (TS) contextual bandit algorithms, in which the algorithm's ability to quantify uncertainty and make decisions depends on the quality of a generative model that is learned offline. Instead of viewing uncertainty in the environment as arising from unobservable latent parameters, our algorithm treats uncertainty as stemming from missing, but potentially observable outcomes (including both future and counterfactual outcomes). If these outcomes were all observed, one could simply make decisions using an "oracle" policy fit on the complete dataset. Inspired by this conceptualization, at each decision-time, our algorithm uses a generative model to probabilistically impute missing outcomes, fits a policy using the imputed complete dataset, and uses that policy to select the next action. We formally show that this algorithm is a generative formulation of TS and establish a state-of-the-art regret bound. Notably, our regret bound depends on the generative model only through the quality of its offline prediction loss, and applies to any method of fitting the "oracle" policy.

强化学习上下文决策生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。