用智能体替代人类指导机器人强化学习,提升训练效率。
Accelerating Robotic Reinforcement Learning with Agent Guidance
- 用多模态智能体充当虚拟教练,提供精准引导
- 在三项任务中样本效率优于人工指导方法
- 适合需要大规模自动化机器人训练的场景
强化学习(RL)为机器人通过试错掌握通用操作技能提供了强大范式,但其实际应用受限于低样本效率。现有基于人类在环(HIL)的方法虽能加速训练,但依赖人类监督导致1:1指导比例,难以扩展,且易因操作员疲劳和能力差异引入高方差。本文提出代理引导策略搜索(AGPS),以多模态智能体取代人类监督者。核心洞察是:该智能体可作为语义世界模型,注入内在价值先验,结构化物理探索过程。通过使用工具,智能体以修正路径点和空间约束的形式提供精确引导,实现探索剪枝。我们在三个任务上验证方法,涵盖精密插入到可变形物体操作。结果表明,AGPS在样本效率上超越现有HIL方法。该框架实现了监督流程自动化,为无劳动力、可扩展的机器人学习开辟道路。项目网站:https://agps-rl.github.io/agps/
原文摘要 · Abstract (English)
Reinforcement Learning (RL) offers a powerful paradigm for autonomous robots to master generalist manipulation skills through trial-and-error. However, its real-world application is stifled by low sample efficiency. Recent Human-in-the-Loop (HIL) methods accelerate training by using human corrections, yet this approach faces a scalability barrier. Reliance on human supervisors imposes a 1:1 supervision ratio that limits scalability, suffers from operator fatigue over extended sessions, and introduces high variance due to inconsistent human proficiency. We present Agent-guided Policy Search (AGPS), a framework that automates the training pipeline by replacing human supervisors with a multimodal agent. Our key insight is that the agent can be viewed as a semantic world model, injecting intrinsic value priors to structure physical exploration. By using tools, the agent provides precise guidance via corrective waypoints and spatial constraints for exploration pruning. We validate our approach on three tasks, ranging from precision insertion to deformable object manipulation. Results demonstrate that AGPS outperforms HIL methods in sample efficiency. This automates the supervision pipeline, unlocking the path to labor-free and scalable robot learning. Project website: https://agps-rl.github.io/agps/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。