arXiv:2603.02008cs.LGcs.AI2026-03被引 1

用时间对比表示引导探索,无需外部奖励就能学会复杂行为。

Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards

  • 基于未来不可预测性,利用时间对比学习构建探索导向的表征。
  • 在运动、操作和具身智能任务中实现了无需奖励的复杂探索行为。
  • 无需距离学习或记忆机制,方法更简洁高效,适合强化学习初学者。

强化学习中的有效探索不仅需要记录已访问位置,还需理解智能体对世界的感知与表征方式。为学习强大表征,智能体应主动探索能提升其环境知识的状态。时间表征可捕捉解决多种潜在任务所需信息,同时避免全状态重建带来的计算开销。本文提出一种利用时间对比表示引导探索的方法,优先选择未来结果不可预测的状态。实验表明,该方法能在运动、操作及具身人工智能任务中实现复杂探索行为,展现出传统需外在奖励才能达成的能力。相比依赖显式距离学习或情景记忆机制(如基于准度量的方法),本方法直接基于时间相似性,提供了一种更简洁但有效的探索策略。

原文摘要 · Abstract (English)

Effective exploration in reinforcement learning requires not only tracking where an agent has been, but also understanding how the agent perceives and represents the world. To learn powerful representations, an agent should actively explore states that contribute to its knowledge of the environment. Temporal representations can capture the information necessary to solve a wide range of potential tasks while avoiding the computational cost associated with full state reconstruction. In this paper, we propose an exploration method that leverages temporal contrastive representations to guide exploration, prioritizing states with unpredictable future outcomes. We demonstrate that such representations can enable the learning of complex exploratory x in locomotion, manipulation, and embodied-AI tasks, revealing capabilities and behaviors that traditionally require extrinsic rewards. Unlike approaches that rely on explicit distance learning or episodic memory mechanisms (e.g., quasimetric-based methods), our method builds directly on temporal similarities, yielding a simpler yet effective strategy for exploration.

强化学习探索策略时间表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。