arXiv:2507.07852cs.LGstat.ML2025-07被引 4

用预训练模型补全缺失数据,理论证明可降低决策误差。

Pre-Trained AI Model Assisted Online Decision-Making under Missing Covariates: A Theoretical Perspective

  • 引入模型弹性概念,衡量真实与补全数据的敏感度差异。
  • 在随机缺失场景下,通过正交学习校准模型,显著提升补全精度。
  • 适合研究预训练模型与决策系统融合的学者参考。

我们研究了一类序列上下文决策问题,其中部分协变量缺失,但可通过预训练AI模型进行填补。从理论角度分析了此类模型对决策过程累积损失(后悔值)的影响。提出新概念‘模型弹性’,量化奖励函数对真实协变量与填补值之间差异的敏感程度,统一刻画因模型填补导致的后悔值,不受缺失机制影响。更令人惊讶的是,在随机缺失(MAR)条件下,可利用正交统计学习和双重稳健回归工具对预训练模型进行序列校准,显著提升填补质量,从而获得更优的后悔界。分析表明,准确的预训练模型在序列决策中具有重要实践价值,且模型弹性或可成为理解并改进预训练模型在各类数据驱动决策中集成的核心指标。

原文摘要 · Abstract (English)

We study a sequential contextual decision-making problem in which certain covariates are missing but can be imputed using a pre-trained AI model. From a theoretical perspective, we analyze how the presence of such a model influences the regret of the decision-making process. We introduce a novel notion called "model elasticity", which quantifies the sensitivity of the reward function to the discrepancy between the true covariate and its imputed counterpart. This concept provides a unified way to characterize the regret incurred due to model imputation, regardless of the underlying missingness mechanism. More surprisingly, we show that under the missing at random (MAR) setting, it is possible to sequentially calibrate the pre-trained model using tools from orthogonal statistical learning and doubly robust regression. This calibration significantly improves the quality of the imputed covariates, leading to much better regret guarantees. Our analysis highlights the practical value of having an accurate pre-trained model in sequential decision-making tasks and suggests that model elasticity may serve as a fundamental metric for understanding and improving the integration of pre-trained models in a wide range of data-driven decision-making problems.

决策优化预训练模型缺失数据理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。