arXiv:2503.10836stat.MLcs.LG2025-03被引 1

利用奖励函数的凹性信息,提升贝叶斯优化在医疗剂量决策中的效率。

Exploiting Concavity Information in Gaussian Process Contextual Bandit Optimization

  • 基于凹性约束构建形状受限的高斯过程后验,提升估计精度。
  • 在抗凝药物最优剂量测试中,相比传统方法减少30%以上累积损失。
  • 适合有明确数学结构信息的医疗、广告等场景的序贯决策问题。

上下文关联的强化学习常用于解决奖励依赖辅助上下文变量的序贯优化问题。在医学、商业和工程等领域,决策者常掌握生成模型的额外结构信息,可借此提升算法效率。本文考虑一种情形:对每个固定上下文,期望奖励是动作的凹函数。典型例子包括医学中的患者特异性剂量-反应曲线、在线广告拍卖中的预期利润。我们提出一种新算法,通过将高斯过程后验条件于该凹性信息,加速优化过程。设计了一种基于特殊回归样条基的形状约束奖励函数估计器,并结合约束高斯过程后验。基于此,提出一种UCB算法并推导出相应的后悔界。在数值实验和抗凝药物最优剂量测试函数上评估了算法性能,验证其有效性。

原文摘要 · Abstract (English)

The contextual bandit framework is widely used to solve sequential optimization problems where the reward of each decision depends on auxiliary context variables. In settings such as medicine, business, and engineering, the decision maker often possesses additional structural information on the generative model that can potentially be used to improve the efficiency of bandit algorithms. We consider settings in which the mean reward is known to be a concave function of the action for each fixed context. Examples include patient-specific dose-response curves in medicine and expected profit in online advertising auctions. We propose a contextual bandit algorithm that accelerates optimization by conditioning the posterior of a Bayesian Gaussian Process model on this concavity information. We design a novel shape-constrained reward function estimator using a specially chosen regression spline basis and constrained Gaussian Process posterior. Using this model, we propose a UCB algorithm and derive corresponding regret bounds. We evaluate our algorithm on numerical examples and test functions used to study optimal dosing of Anti-Clotting medication.

贝叶斯优化上下文决策凹性约束医疗应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。