arXiv:2511.04598cs.LG2025-11

让智能体自主设定目标,无需奖励就能学会通用任务。

Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning

  • 智能体自选目标,环境无关训练,效率接近有引导强化学习。
  • 平均任务成功率提升并趋于稳定,单个目标表现波动但整体可靠。
  • 适合通用预训练,部署前无需特定任务微调。

本文研究将常规强化学习环境转化为目标条件环境,使智能体能在无奖励、自主状态下学会解决任务。实验表明,智能体可通过环境无关的方式自选目标,在训练时间上与外部引导的强化学习相当。该方法不依赖具体离策略学习算法,因不偏爱任何目标,导致单个目标性能不稳定;但平均成功率持续提升并趋于稳定。经此方法训练的智能体可被指令寻找环境中任意观测到的状态,实现任务通用的预训练,适用于后续多样化具体应用。

原文摘要 · Abstract (English)

In this paper we study how transforming regular reinforcement learning environments into goal-conditioned environments can let agents learn to solve tasks autonomously and reward-free. We show that an agent can learn to solve tasks by selecting its own goals in an environment-agnostic way, at training times comparable to externally guided reinforcement learning. Our method is independent of the underlying off-policy learning algorithm. Since our method is environment-agnostic, the agent does not value any goals higher than others, leading to instability in performance for individual goals. However, in our experiments, we show that the average goal success rate improves and stabilizes. An agent trained with this method can be instructed to seek any observations made in the environment, enabling generic training of agents prior to specific use cases.

强化学习自监督目标导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。