arXiv:2510.09908stat.MLcs.LG2025-10

用预训练模型补全缺失上下文,提升在线决策效果

Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation

  • 用预训练模型在在线决策中补全缺失特征
  • 理论证明补全误差影响可量化,性能接近最优
  • 适合有历史完整数据、需处理不完整上下文的场景

大规模预训练模型的兴起使得低成本生成预测或合成特征成为可能,这引出了如何将这些替代预测融入下游决策的问题。本文研究在线线性上下文老虎机场景,其中上下文复杂且非平稳,仅部分可观测。除老虎机数据外,假设存在一个包含完整上下文的辅助数据集——这在实践中常见,因这类数据无需自适应干预即可收集。我们提出 PULSE-UCB 算法,利用在辅助数据上训练的预训练模型,在在线决策中对缺失特征进行插补。我们建立了后悔率上界,其分解为标准老虎机项与反映预训练模型质量的附加项。在独立同分布上下文且缺失特征满足 Hölder 光滑性条件下,PULSE-UCB 达到近最优性能,并有匹配的下界支持。结果量化了预测上下文不确定性对决策质量的影响,以及改善下游学习所需的历史数据量。

原文摘要 · Abstract (English)

The rise of large-scale pretrained models has made it feasible to generate predictive or synthetic features at low cost, raising the question of how to incorporate such surrogate predictions into downstream decision-making. We study this problem in the setting of online linear contextual bandits, where contexts may be complex, nonstationary, and only partially observed. In addition to bandit data, we assume access to an auxiliary dataset containing fully observed contexts--common in practice since such data are collected without adaptive interventions. We propose PULSE-UCB, an algorithm that leverages pretrained models trained on the auxiliary data to impute missing features during online decision-making. We establish regret guarantees that decompose into a standard bandit term plus an additional component reflecting pretrained model quality. In the i.i.d. context case with Hölder-smooth missing features, PULSE-UCB achieves near-optimal performance, supported by matching lower bounds. Our results quantify how uncertainty in predicted contexts affects decision quality and how much historical data is needed to improve downstream learning.

上下文老虎机预训练模型特征补全在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。