研究如何让世界模型被误导,揭示了深度强化学习探索的理论边界。
How Hard is it to Confuse a World Model?
- 用约束优化找最易混淆的模型,使最优与次优策略表现差异大。
- 模型越不确定,越容易被构造出混淆实例,相关性显著。
- 为深度模型基础强化学习提供可理论支撑的探索策略设计依据。
在强化学习理论中,最混淆实例是建立后悔下界的核心概念,即解决问题所需的最小探索量。给定参考模型及其最优策略,最混淆实例是指在统计上与参考模型尽可能接近,却使得次优策略成为最优的替代模型。尽管该概念在多臂赌博机和遍历性表格马尔可夫决策过程中有深入研究,但在一般情形下仍属开放问题。本文将此问题形式化为神经网络世界模型的约束优化:寻找一个与参考模型统计上接近但导致最优与次优策略性能显著分化的修改模型。我们提出一种对抗训练方法求解,并在不同质量的世界模型上进行实证研究。结果表明,可实现的混淆程度与近似模型的不确定性相关,这可能为深度模型基础强化学习提供理论驱动的探索策略。
原文摘要 · Abstract (English)
In reinforcement learning (RL) theory, the concept of most confusing instances is central to establishing regret lower bounds, that is, the minimal exploration needed to solve a problem. Given a reference model and its optimal policy, a most confusing instance is the statistically closest alternative model that makes a suboptimal policy optimal. While this concept is well-studied in multi-armed bandits and ergodic tabular Markov decision processes, constructing such instances remains an open question in the general case. In this paper, we formalize this problem for neural network world models as a constrained optimization: finding a modified model that is statistically close to the reference one, while producing divergent performance between optimal and suboptimal policies. We propose an adversarial training procedure to solve this problem and conduct an empirical study across world models of varying quality. Our results suggest that the degree of achievable confusion correlates with uncertainty in the approximate model, which may inform theoretically-grounded exploration strategies for deep model-based RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。