arXiv:2605.23351cs.LGcs.GT2026-05

提出新算法在延迟反馈下仍能保持安全策略的稳定表现。

Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays

论文配图:Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays
图 1 · 摘自论文原文
  • 设计延迟自适应的重启阈值,动态调整探索与保守策略
  • 实现近似常数的安全基线误差与最优最坏情况后悔上界
  • 适合对安全性要求高的在线决策场景,如金融或医疗

我们研究了带延迟反馈和不带延迟反馈的对抗性多臂赌博机问题,在保证安全性的目标下:在最坏情况下实现极小化后悔率的同时,相对于指定的“安全”基准策略的后悔率几乎恒定。现有方法可在即时反馈下平衡此权衡,但任意延迟可能错配保守与探索的切换时机,危及安全保证。为此,我们提出普鲁登特-银行家(Prudent-Banker)算法,结合延迟自适应的在线镜面下降与改进的分阶段激进机制。其关键技术贡献在于一个延迟校准的重启阈值,严格考虑未观测反馈带来的最坏情况失真,并可靠检测比较器次优性。我们还建立了新的安全约束下对抗性延迟赌博机的下界,证明普鲁登特-银行家的后悔率界在对数因子内不可改进。据我们所知,该算法是首个在有无延迟情况下均实现最优安全-鲁棒性权衡的算法:伪后悔为 $\widetilde{O}(\sqrt{T}+\sqrt{D})$,且对安全比较器的后悔为 $\widetilde{O}(1)$。在多种延迟分布下的实验表明,相比标准延迟鲁棒基线,普鲁登特-银行家能有效平衡安全与学习。

原文摘要 · Abstract (English)

We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while keeping nearly constant regret relative to a designated "safe" baseline policy. Existing approaches can balance this trade-off with immediate feedback for smooth comparators, but arbitrary delays can mistime transitions between conservatism and exploration, endangering the safety guarantee. To bridge this gap, we propose Prudent-Banker, a novel algorithm that combines a delay-adapted variant of Online Mirror Descent with a modified phased-aggression mechanism. Its key technical contribution is a delay-calibrated restart threshold that rigorously accounts for the worst-case distortion induced by unobserved feedback and reliably detects comparator suboptimality. We also establish new lower bounds for safety-constrained adversarial delayed bandits, showing that the regret guarantees of Prudent-Banker are unimprovable, up to logarithmic factors, under the baseline-safety requirement. To the best of our knowledge, Prudent-Banker is the first algorithm to achieve the optimal safety--robustness trade-off: pseudo-regret $\widetilde{O}(\sqrt{T}+\sqrt{D})$ together with $\widetilde{O}(1)$ regret against the safe comparator, both with and without delays. Experiments across diverse delay distributions show that, unlike standard delay-robust baselines, Prudent-Banker effectively balances safety and learning.

强化学习在线学习安全决策延迟反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。