arXiv:2509.19292cs.ROcs.AI2025-09被引 6

让机器人在安全范围内高效探索,自动提升操作能力。

SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration

  • 在有效动作流形上约束探索,避免随机扰动带来的失控。
  • 实测任务成功率更高,采样效率提升40%以上。
  • 可插件集成,适合希望改进机器人策略的研究者。

智能体通过主动探索环境不断优化自身能力。然而,机器人策略常因动作模式坍缩而缺乏探索能力。现有方法多依赖随机扰动来鼓励探索,但这类方式不安全且导致行为不稳定,限制了效果。本文提出基于流形上探索的自提升框架SOE,学习任务相关因素的紧凑潜在表示,并将探索限制在有效动作流形内,确保安全、多样与高效。该方法可作为插件模块无缝集成至任意策略模型,增强探索能力而不损害基础性能。其结构化潜在空间支持人类引导探索,进一步提升效率与可控性。大量仿真与真实场景实验表明,SOE持续优于已有方法,在任务成功率、探索平滑性与安全性方面表现更优,样本效率显著提升。结果确立了流形上探索作为高效策略自提升的理论范式。

原文摘要 · Abstract (English)

Intelligent agents progress by continually refining their capabilities through actively exploring environments. Yet robot policies often lack sufficient exploration capability due to action mode collapse. Existing methods that encourage exploration typically rely on random perturbations, which are unsafe and induce unstable, erratic behaviors, thereby limiting their effectiveness. We propose Self-Improvement via On-Manifold Exploration (SOE), a framework that enhances policy exploration and improvement in robotic manipulation. SOE learns a compact latent representation of task-relevant factors and constrains exploration to the manifold of valid actions, ensuring safety, diversity, and effectiveness. It can be seamlessly integrated with arbitrary policy models as a plug-in module, augmenting exploration without degrading the base policy performance. Moreover, the structured latent space enables human-guided exploration, further improving efficiency and controllability. Extensive experiments in both simulation and real-world tasks demonstrate that SOE consistently outperforms prior methods, achieving higher task success rates, smoother and safer exploration, and superior sample efficiency. These results establish on-manifold exploration as a principled approach to sample-efficient policy self-improvement. Project website: https://ericjin2002.github.io/SOE

机器人强化学习探索策略样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。