用不完整演示数据引导强化学习探索,提升复杂环境下的适应能力。
Mixture of Autoencoder Experts Guidance using Unlabeled and Incomplete Data for Exploration in Reinforcement Learning
- 通过自编码器专家混合模型捕捉多样行为,处理缺失演示数据。
- 在稀疏与密集奖励环境中均实现稳定探索与高效学习性能。
- 适合真实场景中数据不全、需精准控制探索的强化学习任务。
强化学习近期趋势强调代理需从无奖励交互和非显式监督信号(如未标记或不完整的示范)中学习,而非仅依赖显式奖励最大化。开发能高效适应现实环境的通用代理常需利用这些无奖励信号来指导学习与行为。然而,尽管内在动机技术可帮助代理在缺乏显式奖励时探索新颖或不确定状态,但在密集奖励环境或高维状态/动作空间下仍面临挑战。此外,多数现有方法直接使用未经处理的内在奖励信号,难以有效塑造或控制代理的探索行为。我们提出一种框架,可有效利用不完整且不完美的专家示范。通过映射函数将代理状态与专家数据的相似性转化为可调控的内在奖励,使探索更灵活、目标明确。采用自编码器专家混合模型以捕捉多样化行为并处理示范中的缺失信息。实验表明,该方法在稀疏与密集奖励环境下均实现稳健探索和优异性能,即使示范数据稀疏或不完整。为实际应用中最优数据不可得且需精确奖励控制的强化学习提供了可行方案。
原文摘要 · Abstract (English)
Recent trends in Reinforcement Learning (RL) highlight the need for agents to learn from reward-free interactions and alternative supervision signals, such as unlabeled or incomplete demonstrations, rather than relying solely on explicit reward maximization. Additionally, developing generalist agents that can adapt efficiently in real-world environments often requires leveraging these reward-free signals to guide learning and behavior. However, while intrinsic motivation techniques provide a means for agents to seek out novel or uncertain states in the absence of explicit rewards, they are often challenged by dense reward environments or the complexity of high-dimensional state and action spaces. Furthermore, most existing approaches rely directly on the unprocessed intrinsic reward signals, which can make it difficult to shape or control the agent's exploration effectively. We propose a framework that can effectively utilize expert demonstrations, even when they are incomplete and imperfect. By applying a mapping function to transform the similarity between an agent's state and expert data into a shaped intrinsic reward, our method allows for flexible and targeted exploration of expert-like behaviors. We employ a Mixture of Autoencoder Experts to capture a diverse range of behaviors and accommodate missing information in demonstrations. Experiments show our approach enables robust exploration and strong performance in both sparse and dense reward environments, even when demonstrations are sparse or incomplete. This provides a practical framework for RL in realistic settings where optimal data is unavailable and precise reward control is needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。