arXiv:2505.00398cs.LG2025-05中稿 · AISTATS, 2025被引 1

首次实现在线优化中全程零约束违规,保障动态变化下的绝对安全。

Safety in the Face of Adversity: Achieving Zero Constraint Violation in Online Learning with Slowly Changing Constraints

  • 采用原对偶方法与对偶空间梯度上升,应对缓慢变化的约束。
  • 通过双参数学习率设计,实现零约束违规与次线性遗憾。
  • 适合对安全性要求极高的实时决策场景,如自动驾驶、医疗系统。

我们首次为在线凸优化(OCO)在所有轮次中提供零约束违规的理论保证,解决了动态约束变化问题。与现有受限在线凸优化方法允许偶尔安全突破不同,本工作在约束逐轮变化量较小的假设下,首次实现了严格安全。方法基于原对偶框架,并在对偶空间中采用在线梯度上升。通过引入分段学习率策略,同时确保零约束违规和次线性遗憾。该框架标志着在面对变化约束时,首次实现可证明的绝对安全性,突破了以往研究的局限。

原文摘要 · Abstract (English)

We present the first theoretical guarantees for zero constraint violation in Online Convex Optimization (OCO) across all rounds, addressing dynamic constraint changes. Unlike existing approaches in constrained OCO, which allow for occasional safety breaches, we provide the first approach for maintaining strict safety under the assumption of gradually evolving constraints, namely the constraints change at most by a small amount between consecutive rounds. This is achieved through a primal-dual approach and Online Gradient Ascent in the dual space. We show that employing a dichotomous learning rate enables ensuring both safety, via zero constraint violation, and sublinear regret. Our framework marks a departure from previous work by providing the first provable guarantees for maintaining absolute safety in the face of changing constraints in OCO.

在线学习安全优化约束满足

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。