arXiv:2511.07730cs.LGcs.RO2025-11被引 3

用多步蒙特卡洛法学习目标间距离,让机器人在长序列任务中更高效地规划路径。

Multistep Quasimetric Learning for Scalable Goal-conditioned Reinforcement Learning

  • 基于多步蒙特卡洛回报构建拟度量距离函数
  • 在4000步的长程模拟任务中超越现有方法
  • 首次实现真实机器人视觉数据的端到端多步拼接

在环境中学习如何达成目标是人工智能的长期挑战,而长时序推理仍是现代方法的难点。核心问题在于如何估计观测对之间的时序距离。尽管时序差分方法通过局部更新提供最优性保证,但其性能常不如采用全局更新(如多步回报)的蒙特卡洛方法,后者缺乏此类保证。本文提出一种实用的离线目标条件强化学习(GCRL)方法,利用多步蒙特卡洛回报拟合一个拟度量距离。实验表明,该方法在长达4000步的模拟任务中,即使使用视觉观测,仍优于现有离线GCRL方法。同时,我们在真实世界机器人操作场景(Bridge设置)中验证了该方法的有效性,首次实现了从无标签视觉观测离线数据集出发的端到端多步拼接,并展现出鲁棒的长程泛化能力。

原文摘要 · Abstract (English)

Learning how to reach goals in an environment is a longstanding challenge in AI, yet reasoning over long horizons remains a challenge for modern methods. The key question is how to estimate the temporal distance between pairs of observations. While temporal difference methods leverage local updates to provide optimality guarantees, they often perform worse than Monte Carlo methods that perform global updates (e.g., with multi-step returns), which lack such guarantees. We show how these approaches can be integrated into a practical offline GCRL method that fits a quasimetric distance using a multistep Monte-Carlo return. We show our method outperforms existing offline GCRL methods on long-horizon simulated tasks with up to 4000 steps, even with visual observations. We also demonstrate that our method can enable stitching in the real-world robotic manipulation domain (Bridge setup). Our approach is the first end-to-end offline GCRL method that enables multistep stitching in this real-world manipulation domain from an unlabeled offline dataset of visual observations and demonstrate robust horizon generalization.

强化学习目标条件长程规划机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。