arXiv:2502.13200cs.LGcs.AI2025-02被引 1

用自监督学习构建世界模型,让智能体自主探索复杂环境。

Learning To Explore With Predictive World Model Via Self-Supervised Learning

  • 通过自监督学习构建内部世界模型,实现无需外部奖励的自主探索
  • 在18个Atari游戏上表现优于现有方法,尤其在稀疏奖励场景
  • 适合研究自主学习、内在动机与强化学习的科研人员

自主人工智能代理必须能在复杂环境中学习行为,而无需人类设计任务和奖励。为每种环境单独设计这些功能不可行,因此推动了内在奖励函数的发展。本文提出利用长期被忽视的认知元素,构建一个内在动机代理的内部世界模型。该代理能与环境进行有效迭代,无需预先设计的奖励函数即可学习复杂行为。我们使用18个Atari游戏评估了需要反应性与规划性行为的游戏中的认知技能涌现情况。结果表明,在密集奖励和稀疏奖励等多种测试场景中,性能均显著优于当前最先进方法。

原文摘要 · Abstract (English)

Autonomous artificial agents must be able to learn behaviors in complex environments without humans to design tasks and rewards. Designing these functions for each environment is not feasible, thus, motivating the development of intrinsic reward functions. In this paper, we propose using several cognitive elements that have been neglected for a long time to build an internal world model for an intrinsically motivated agent. Our agent performs satisfactory iterations with the environment, learning complex behaviors without needing previously designed reward functions. We used 18 Atari games to evaluate what cognitive skills emerge in games that require reactive and deliberative behaviors. Our results show superior performance compared to the state-of-the-art in many test cases with dense and sparse rewards.

自主学习内在动机世界模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。