针对带扰动的近线性函数,提出对抗性强化学习优化新方法
Adversarial bandit optimization for approximately linear functions
- 在每轮中损失函数为线性+任意扰动,基于观测选择调整策略
- 给出期望与高概率下的后悔界,改进线性情形的已有结果
- 适用于非凸非光滑场景,适合研究在线优化的学者
我们研究非凸、非光滑函数的对抗性强化学习优化问题,其中每轮的损失函数由线性部分和在观察到玩家选择后任意选定的小扰动组成。本文给出了该问题的期望后悔界与高概率后悔界。我们的结果还表明,在无扰动的对抗性线性优化这一特例中,可获得更优的高概率后悔界。同时,我们还给出了期望后悔的下界。
原文摘要 · Abstract (English)
We consider a bandit optimization problem for nonconvex and non-smooth functions, where in each trial the loss function is the sum of a linear function and a small but arbitrary perturbation chosen after observing the player's choice. We give both expected and high probability regret bounds for the problem. Our result also implies an improved high-probability regret bound for the bandit linear optimization, a special case with no perturbation. We also give a lower bound on the expected regret.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。