揭示无监督强化学习中自发探索的内在机制,发现其靠低秩表征驱动。
Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
- 通过对比学习构建隐式奖励,自动调节探索与利用
- 在长时序目标任务中实现无需外部奖励的自主探索
- 适用于需要安全探索的复杂环境,如机器人导航
本文首次深入探究无监督强化学习中自发探索的机制。研究单目标对比强化学习(SGCRL),一种无需外部奖励或教学课程即可解决高难度长时序目标到达任务的自监督算法。结合理论分析与受控实验,发现SGCRL通过学习的状态表示生成隐式奖励,自动调整奖励结构:在抵达目标前促进探索,达成后转向利用。实验表明,这种探索行为源于对状态空间的低秩表示学习,而非神经网络函数逼近。该理解使我们能够改进SGCRL以实现安全感知探索。
原文摘要 · Abstract (English)
In this work, we take a first step toward elucidating the mechanisms behind emergent exploration in unsupervised reinforcement learning. We study Single-Goal Contrastive Reinforcement Learning (SGCRL), a self-supervised algorithm capable of solving challenging long-horizon goal-reaching tasks without external rewards or curricula. We combine theoretical analysis of the algorithm's objective function with controlled experiments to understand what drives its exploration. We show that SGCRL maximizes implicit rewards shaped by its learned representations. These representations automatically modify the reward landscape to promote exploration before reaching the goal and exploitation thereafter. Our experiments also demonstrate that these exploration dynamics arise from learning low-rank representations of the state space rather than from neural network function approximation. Our improved understanding enables us to adapt SGCRL to perform safety-aware exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。