用随机特征近似信息增益,实现可解释的深度强化学习探索。
Information-Based Exploration via Random Features for Reinforcement Learning
- 基于随机傅里叶特征近似信息增益,避免神经网络黑箱不确定性估计。
- 在多个控制与导航任务中性能媲美主流探索方法,理论误差有界。
- 适合关注探索机制可解释性与理论保障的研究者与工程师。
表示学习使经典探索策略得以拓展至深度强化学习,但常导致算法复杂且理论分析困难。本文提出基于贝叶斯核方法理论的随机特征信息增益(RFIG),利用随机傅里叶特征近似信息增益,并在不可数空间中计算探索奖励。我们给出了信息增益近似的误差界,避免了基于神经网络的不确定性估计的黑箱特性,支持基于乐观主义的探索。文中提供了实用细节,使RFIG可扩展至深度强化学习场景,能平滑集成到标准深度强化学习算法中。在多种控制与导航任务上的实验表明,RFIG在性能上可与成熟深度探索方法竞争,同时具备更优的理论可解释性。
原文摘要 · Abstract (English)
Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical guarantees harder to establish. We introduce Random Feature Information Gain (RFIG), grounded in Bayesian kernel methods theory, which uses random Fourier features to approximate information gain and compute exploration bonuses in non-countable spaces. We provide error bounds on information gain approximation and avoid the black-box aspects of neural network-based uncertainty estimation, for optimism-based exploration. We present practical details that make RFIG scalable to deep RL scenarios, enabling smooth integration into standard deep RL algorithms. Experimental evaluation across diverse control and navigation tasks demonstrates that RFIG achieves competitive performance with well-established deep exploration methods while offering superior theoretical interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。