提出新方法在连续动作中实时约束成本,提升决策安全性。
Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions
- 用实时成本约束替代期望成本,更适应波动性环境。
- 理论证明误差随时间增长呈根号关系,且实验显著降低违规率。
- 适合医疗剂量、自动驾驶等高风险决策场景使用。
上下文老虎机是不确定性下序列决策的标准框架,应用于临床试验、剂量选择、推荐系统和自主系统。安全在这些应用中至关重要,因为剂量选择或自动驾驶中的单次不安全决策可能带来灾难性后果。现有方法通常为每个动作定义奖励和成本信号,并在成本期望低于阈值的条件下优化奖励。然而,在异方差环境下,选择的动作不仅影响期望收益与成本,还影响观测结果的变异性。本文研究一维连续动作的上下文老虎机,采用阶段性的高概率实现实时成本约束。提出高概率约束UCB算法,通过乐观探索收益、保守估计安全动作集来实现安全。对于线性收益与成本模型,证明了紧致的$ ilde{ ext{O}}(d oot{T})$ regret界,并将分析扩展至一般函数类,基于弹射维度。实验表明,与基于期望成本的基线相比,实现实时成本约束可显著减少违规情况。
原文摘要 · Abstract (English)
Contextual bandits are a standard framework for sequential decision-making under uncertainty, with applications in clinical trials, dosage selection, recommendation systems, and autonomous systems. Safety is central in many of these applications, since a single unsafe decision in settings such as dosage selection or autonomous driving can have catastrophic consequences. A common way to model safety in bandit problems is to associate each action with both a reward signal and a cost signal, and to optimize reward subject to constraints on cost. Most existing safety-constrained bandit models enforce safety by requiring the expected cost of each action to remain below a prescribed threshold. However, this may be insufficient in heteroscedastic settings, where the chosen action affects not only the expected reward and cost, but also the variability of the observed outcomes. We study contextual bandits with one-dimensional continuous actions and stage-wise high-probability constraints on the realized cost. We propose High-Probability Constrained UCB, an optimistic-pessimistic algorithm that explores for reward while conservatively estimating the safe action set. For linear reward and cost models, we prove a tight $\tilde{\mathcal{O}}(d\sqrt{T})$ regret bound, and we extend the analysis to general function classes using the eluder dimension. Experiments show that enforcing realized-cost safety substantially reduces violations compared with expected-cost constrained baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。