在奖励不确定时,用概率保证方法优化个性化治疗方案。
PAC-Bayesian Reward-Certified Outcome Weighted Learning
- 基于概率保证构建保守奖励,使学习更稳健。
- 理论证明可将鲁棒策略学习转化为分类任务,实现有限样本保证。
- 自动校准与高效优化,适合高风险决策场景应用。
通过结果加权学习(OWL)估计最优个体化治疗规则(ITRs)通常依赖于噪声或乐观的观测奖励,而这些奖励未必反映真实的潜在效用。忽略奖励不确定性会导致选择表现虚高的策略,但现有OWL框架缺乏有限样本保证来系统地将不确定性纳入学习目标。为此,我们提出PAC-Bayesian Reward-Certified Outcome Weighted Learning(PROWL)。在单边不确定性证书下,PROWL构造一个保守奖励和严格依赖策略的真期望值下界。理论上,我们证明了一个精确的认证缩减,将鲁棒策略学习转化为统一、无需分裂的成本敏感分类任务。该形式化允许推导随机ITR的非渐近PAC-Bayes下界,其中我们证明最大化该下界的最优后验恰好由广义贝叶斯更新刻画。为克服广义贝叶斯推断中学习率选择问题,我们引入完全自动化的基于边界校准程序,并结合费希尔一致的认证铰链损失以实现高效优化。实验表明,在严重奖励不确定性下,PROWL相较于标准ITR估计方法能更准确地识别稳健且高价值的治疗方案。
原文摘要 · Abstract (English)
Estimating optimal individualized treatment rules (ITRs) via outcome weighted learning (OWL) often relies on observed rewards that are noisy or optimistic proxies for the true latent utility. Ignoring this reward uncertainty leads to the selection of policies with inflated apparent performance, yet existing OWL frameworks lack the finite-sample guarantees required to systematically embed such uncertainty into the learning objective. To address this issue, we propose PAC-Bayesian Reward-Certified Outcome Weighted Learning (PROWL). Given a one-sided uncertainty certificate, PROWL constructs a conservative reward and a strictly policy-dependent lower bound on the true expected value. Theoretically, we prove an exact certified reduction that transforms robust policy learning into a unified, split-free cost-sensitive classification task. This formulation enables the derivation of a nonasymptotic PAC-Bayes lower bound for randomized ITRs, where we establish that the optimal posterior maximizing this bound is exactly characterized by a general Bayes update. To overcome the learning-rate selection problem inherent in generalized Bayesian inference, we introduce a fully automated, bounds-based calibration procedure, coupled with a Fisher-consistent certified hinge surrogate for efficient optimization. Our experiments demonstrate that PROWL achieves improvements in estimating robust, high-value treatment regimes under severe reward uncertainty compared to standard methods for ITR estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。