arXiv:2603.26441cs.RO2026-03

120分钟内用笔记本完成真实机器人图像目标导航训练。

120 Minutes and a Laptop: Minimalist Image-goal Navigation via Unsupervised Exploration and Offline RL

  • 无监督探索+事后目标重标注,离线强化学习实现高效数据收集
  • 在仿真与真实环境均超越零样本基线,数据量越大效果越佳
  • 全程无需人工干预,适合快速原型开发与部署

当前图像目标视觉导航多依赖大规模数据集、大量预训练和高算力资源。本文提出MINav方法,证明可在120分钟内、仅用消费级笔记本、完全无需人工干预的条件下,完成数据采集、域内策略训练并部署至真实世界。该方法将图像目标导航建模为离线目标条件强化学习问题,结合无监督数据收集、事后目标重标注与离线策略学习。仿真与真实世界实验表明,MINav提升探索效率,在目标环境中优于零样本导航基线,且随数据集规模增长表现持续优化。结果表明,高计算效率下仍可实现有效真实机器人学习,显著降低快速策略原型化与部署的门槛。

原文摘要 · Abstract (English)

The prevailing paradigm for image-goal visual navigation often assumes access to large-scale datasets, substantial pretraining, and significant computational resources. In this work, we challenge this assumption. We show that we can collect a dataset, train an in-domain policy, and deploy it to the real world (1) in less than 120 minutes, (2) on a consumer laptop, (3) without any human intervention. Our method, MINav, formulates image-goal navigation as an offline goal-conditioned reinforcement learning problem, combining unsupervised data collection with hindsight goal relabeling and offline policy learning. Experiments in simulation and the real world show that MINav improves exploration efficiency, outperforms zero-shot navigation baselines in target environments, and scales favorably with dataset size. These results suggest that effective real-world robotic learning can be achieved with high computational efficiency, lowering the barrier to rapid policy prototyping and deployment.

机器人学习离线强化学习无监督探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。