未训练的神经网络策略会自发产生有结构的探索行为。
Exploration Behavior of Untrained Policies
- 用无限宽网络理论分析未训练策略的初始探索模式
- 发现初始策略能生成非随机的状态访问分布
- 为早期训练探索提供可设计的初始化方法
探索仍是强化学习中的核心挑战,尤其在奖励稀疏或对抗性环境中。本文研究深度神经网络策略在训练前如何通过架构隐式塑造探索行为。通过理论与实证分析,在简化模型中展示未训练策略可生成弹道式或扩散式轨迹。基于无限宽网络理论和连续时间极限,发现未训练策略会产生相关动作,导致非平凡的状态访问分布。针对标准架构的轨迹分布分析揭示了探索的归纳偏置。研究成果建立了理论与实验框架,将策略初始化作为理解早期训练探索行为的设计工具。
原文摘要 · Abstract (English)
Exploration remains a fundamental challenge in reinforcement learning (RL), particularly in environments with sparse or adversarial reward structures. In this work, we study how the architecture of deep neural policies implicitly shapes exploration before training. We theoretically and empirically demonstrate strategies for generating ballistic or diffusive trajectories from untrained policies in a toy model. Using the theory of infinite-width networks and a continuous-time limit, we show that untrained policies return correlated actions and result in non-trivial state-visitation distributions. We discuss the distributions of the corresponding trajectories for a standard architecture, revealing insights into inductive biases for tackling exploration. Our results establish a theoretical and experimental framework for using policy initialization as a design tool to understand exploration behavior in early training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。