arXiv:2411.06329cs.LGstat.ML2024-11被引 3

高维在线决策中,新方法兼顾后悔最小化与统计推断,突破传统权衡。

Regret Minimization and Statistical Inference in Online Decision Making with High-dimensional Covariates

  • 结合ε-贪婪与硬阈值法,用反倾向加权去偏提升参数估计
  • 在多样协变量下,纯贪婪算法可实现对数级后悔与一致推断
  • 适用于医疗剂量等高维在线决策场景,支持有效政策价值推断

本文研究基于稀疏线性上下文老虎机模型的高维在线决策中的后悔最小化与统计推断及其相互关系。将ε-贪婪老虎机算法与硬阈值法结合用于稀疏参数估计,并引入基于逆倾向加权的去偏推断框架。在边界条件成立时,该方法可达到O(T^{1/2})的后悔或经典的O(T^{1/2})一致性推断,表明探索与利用之间存在不可避免的权衡。若满足协变量多样性条件,我们证明纯贪婪算法(无探索)配合基于平均加权的去偏估计器可同时实现最优的O(log T)后悔和O(T^{1/2})一致性推断。此外,简单样本均值估计器亦可为最优策略价值提供有效推断。数值模拟及华法林用药数据实验验证了方法有效性。

原文摘要 · Abstract (English)

This paper investigates regret minimization, statistical inference, and their interplay in high-dimensional online decision-making based on the sparse linear context bandit model. We integrate the $\varepsilon$-greedy bandit algorithm for decision-making with a hard thresholding algorithm for estimating sparse bandit parameters and introduce an inference framework based on a debiasing method using inverse propensity weighting. Under a margin condition, our method achieves either $O(T^{1/2})$ regret or classical $O(T^{1/2})$-consistent inference, indicating an unavoidable trade-off between exploration and exploitation. If a diverse covariate condition holds, we demonstrate that a pure-greedy bandit algorithm, i.e., exploration-free, combined with a debiased estimator based on average weighting can simultaneously achieve optimal $O(\log T)$ regret and $O(T^{1/2})$-consistent inference. We also show that a simple sample mean estimator can provide valid inference for the optimal policy's value. Numerical simulations and experiments on Warfarin dosing data validate the effectiveness of our methods.

在线决策高维数据统计推断后悔最小化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。