让智能体学会在任意状态间导航,提升世界模型的泛化能力。
Learning World Models for Unconstrained Goal Navigation
- 设计新算法MUN,可建模重放缓冲区中任意子目标间的转移。
- 实验显示世界模型可靠性提升,策略在新目标下泛化能力显著增强。
- 适合研究目标导向强化学习与世界模型泛化的研究人员。
学习世界模型为稀疏奖励下的目标条件强化学习提供了有前景的路径。通过允许智能体在不直接与环境交互的情况下规划动作或探索性目标,世界模型提升了探索效率。世界模型的质量取决于智能体重放缓冲区中数据的丰富程度,期望其在记录轨迹周围的状态空间中具备合理泛化能力。然而,在沿记录轨迹反向或跨不同轨迹的状态间泛化时,学习到的世界模型面临挑战,限制了其对真实世界动态的准确建模。为解决这些问题,我们提出一种新型目标导向探索算法MUN(全称:用于无约束目标导航的世界模型)。该算法能够建模重放缓冲区中任意子目标状态之间的状态转移,从而促进学习可在任意“关键”状态间导航的策略。实验结果表明,MUN增强了世界模型的可靠性,并显著提升了策略在新目标设置下的泛化能力。
原文摘要 · Abstract (English)
Learning world models offers a promising avenue for goal-conditioned reinforcement learning with sparse rewards. By allowing agents to plan actions or exploratory goals without direct interaction with the environment, world models enhance exploration efficiency. The quality of a world model hinges on the richness of data stored in the agent's replay buffer, with expectations of reasonable generalization across the state space surrounding recorded trajectories. However, challenges arise in generalizing learned world models to state transitions backward along recorded trajectories or between states across different trajectories, hindering their ability to accurately model real-world dynamics. To address these challenges, we introduce a novel goal-directed exploration algorithm, MUN (short for "World Models for Unconstrained Goal Navigation"). This algorithm is capable of modeling state transitions between arbitrary subgoal states in the replay buffer, thereby facilitating the learning of policies to navigate between any "key" states. Experimental results demonstrate that MUN strengthens the reliability of world models and significantly improves the policy's capacity to generalize across new goal settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。