首个无需环境模型的通用强化学习智能体,理论证明渐近最优。
A Model-Free Universal AI
- 不依赖环境模型,直接对动作价值分布进行泛化归纳
- 在真实世界假设下,渐近达到ε最优,且为ε贝叶斯最优
- 为通用智能体研究提供新范式,适合理论与强化学习研究者
在一般强化学习中,所有已知最优智能体(包括AIXI)均为模型驱动型,需显式维护并使用环境模型。本文提出首个理论上证明渐近ε-最优的无模型智能体:基于Q-归纳的通用人工智能(AIQI)。AIQI对分布化的动作价值函数进行泛化归纳,而非传统方法中的策略或环境模型。在‘真实世界’假设条件下,我们证明了AIQI具有强渐近ε-最优性及渐近ε-贝叶斯最优性。同时,利用新证明技术,我们亦展示了自洽AIXI在无额外假设下的渐近ε-最优性。研究显著拓展了通用智能体的种类。
原文摘要 · Abstract (English)
In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the first model-free agent proven to be asymptotically $\varepsilon$-optimal in general RL. AIQI performs universal induction over distributional action-value functions, instead of policies or environments like previous works. Under a grain of truth condition, we prove that AIQI is strong asymptotically $\varepsilon$-optimal and asymptotically $\varepsilon$-Bayes-optimal. We also apply our novel proof techniques to show asymptotic $\varepsilon$-optimality of Self-AIXI without any ad-hoc assumptions. Our results significantly expand the diversity of known universal agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。