arXiv:2606.19891cs.LG2026-06

在非凸损失下,用全局扰动预算控制对抗性干扰的强化学习优化方法。

Adversarial Bandit Optimization with Globally Bounded Perturbations to Convex Losses

  • 设计新算法,应对观测后选择的对抗扰动。
  • 首次给出带扰动预算时的期望后悔上界,明确扰动影响。
  • 适用于有噪声或恶意干扰的在线学习场景,适合鲁棒优化研究者。

我们研究非凸、非光滑损失下的对抗性老虎机优化问题。每轮中,学习者选择动作并仅观测该动作的损失。损失由一个基础凸且β-光滑成分和一个可在观察到学习者动作后选择的对抗性扰动组成。扰动受全局预算约束,即其累积幅度受限。该框架将先前针对线性损失的全局预算、动作后扰动模型扩展至一般凸和β-光滑损失。针对这一更广泛类别,我们建立了显式刻画扰动预算影响的期望后悔上界。为实现此目标,我们改进了标准老虎机优化算法,并开发了分析方法以控制扰动带来的额外后悔。在无扰动情况下,结果退化为标准老虎机凸优化设置中β-光滑损失的后悔保证。

原文摘要 · Abstract (English)

We study adversarial bandit optimization in which the loss functions may be non-convex and non-smooth. In each round, the learner selects an action and observes only the loss incurred at that action. The loss consists of an underlying convex and $β$-smooth component and an adversarial perturbation that may be chosen after observing the learner's action. The perturbations are subject to a global budget controlling their cumulative magnitude over time. This framework extends the globally budgeted, post-action perturbation model from underlying linear losses to general convex and $β$-smooth losses. For this broader class, we establish expected regret guarantees that explicitly characterize the effect of the perturbation budget. To establish these guarantees, we modify a standard bandit optimization algorithm and develop an analysis that controls the additional regret caused by the perturbations. In the absence of perturbations, our results reduce to regret guarantees for the standard bandit convex optimization setting with $β$-smooth losses.

强化学习对抗优化后悔分析凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。