arXiv:2411.01396cs.LGcs.AI2024-11NeurIPS被引 1

让智能体优先探索能到达的边界目标,提升未知环境探索效率。

Exploring the Edges of Latent State Clusters for Goal-Conditioned Reinforcement Learning

  • 在潜在空间聚类可达状态,优先选边界上可触及的目标
  • 在复杂机器人任务中探索速度显著优于基线方法
  • 适合需要高效自主探索的强化学习场景

在无监督的目标条件强化学习中,高效探索未知环境是核心挑战。尽管选择先前探索区域边界的探索目标是一种有效策略,但训练中的策略可能仍难以抵达边界上的稀有目标,导致探索能力下降。我们提出“聚类边缘探索”(CE²),一种新的目标导向探索算法:在选择稀疏区域的目标时,优先考虑当前策略仍可到达的状态。其核心思想是在潜在空间中对由当前策略易达的状态进行聚类,并在执行探索前,优先遍历这些聚类边界上具有高探索潜力的状态。在多足蚂蚁机器人走迷宫、机械臂在杂乱桌面上操作物体、仿人机械手抓握旋转物体等复杂机器人环境中,CE² 相较于基线方法和消融实验,展现出更优的探索效率。

原文摘要 · Abstract (English)

Exploring unknown environments efficiently is a fundamental challenge in unsupervised goal-conditioned reinforcement learning. While selecting exploratory goals at the frontier of previously explored states is an effective strategy, the policy during training may still have limited capability of reaching rare goals on the frontier, resulting in reduced exploratory behavior. We propose "Cluster Edge Exploration" ($CE^2$), a new goal-directed exploration algorithm that when choosing goals in sparsely explored areas of the state space gives priority to goal states that remain accessible to the agent. The key idea is clustering to group states that are easily reachable from one another by the current policy under training in a latent space and traversing to states holding significant exploration potential on the boundary of these clusters before doing exploratory behavior. In challenging robotics environments including navigating a maze with a multi-legged ant robot, manipulating objects with a robot arm on a cluttered tabletop, and rotating objects in the palm of an anthropomorphic robotic hand, $CE^2$ demonstrates superior efficiency in exploration compared to baseline methods and ablations.

强化学习探索策略机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。