提出鲁棒在线决策框架,放宽了传统强化学习的假设。
Regret Bounds for Robust Online Decision Making
- 用凸集表示决策对应的结果分布,允许对抗性选择
- 在鲁棒线性老虎机和表格型强化学习中实现更优的后悔界
- 适合关注模型不确定性与对抗环境的研究者
我们提出一个框架,通过允许鲁棒(多值)模型来推广“具有结构化观测的决策”。在此框架中,每个模型将每个决策与结果上的一组凸概率分布相关联。自然可以任意(对抗性地)从该集合中选择分布,且可依赖历史信息。该框架比经典老虎机和强化学习更具一般性,因为可实现性假设变得更弱且更现实。随后我们推导了该框架下的后悔界理论。尽管上下界不紧,但足以完全刻画幂律可学习性。我们在两个特例中验证该理论:鲁棒线性老虎机和表格型鲁棒在线强化学习。在这两种情况下,我们都得到了优于现有最优的后悔界(未考虑计算效率)。
原文摘要 · Abstract (English)
We propose a framework which generalizes "decision making with structured observations" by allowing robust (i.e. multivalued) models. In this framework, each model associates each decision with a convex set of probability distributions over outcomes. Nature can choose distributions out of this set in an arbitrary (adversarial) manner, that can be nonoblivious and depend on past history. The resulting framework offers much greater generality than classical bandits and reinforcement learning, since the realizability assumption becomes much weaker and more realistic. We then derive a theory of regret bounds for this framework. Although our lower and upper bounds are not tight, they are sufficient to fully characterize power-law learnability. We demonstrate this theory in two special cases: robust linear bandits and tabular robust online reinforcement learning. In both cases, we derive regret bounds that improve state-of-the-art (except that we do not address computational efficiency).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。