用简单任务的反向轨迹训练复杂操作,降低人工示范成本。
Reverse to Advance: Teleoperation-Cost Effective Hard Policy Learning from Reversed Easy Tasks

- 通过反转易任务生成可复用数据,自动收集并迭代优化
- 在模拟与真实机器人上实现成功率提升,数据效率更高
- 适合需高精度操作但难获取示范的场景
高质量遥操作数据采集成本高昂,尤其针对困难任务。我们观察到许多任务存在方向不对称性:完成正向困难任务困难,而通过放松或破坏环境反转任务则相对容易。这表明反转后的易任务轨迹可作为困难任务的可扩展监督信号,从而降低人工示范采集成本。然而,反转数据可能噪声大,直接训练效果不佳。为此,我们提出一种基于时间反转易任务的低成本遥操作框架,包含三个核心组件:闭环数据收集流程,交替使用困难与易任务策略自动重置环境并生成多样化轨迹;分层数据精炼流程,对易任务轨迹进行时间反转,并利用运动学先验和批评者引导的优势过滤低质量动作;迭代策略学习方法,持续在线学习中结合初始反转数据与过滤后的反转数据训练困难任务策略。该方法实现了复杂高精度操作任务的可扩展、可靠训练。在两个仿真基准及真实机器人实验中,相比基于反转和强化学习的基线,本方法以更高效的数据使用和更稳定的训练提升困难任务成功率,且无需大量困难任务的遥操作示范。
原文摘要 · Abstract (English)
High-quality teleoperation datasets are costly to collect, particularly for hard tasks. We observe that many tasks exhibit directional asymmetry: completing the forward hard task is difficult, whereas reversing it by relaxing or disrupting the environment is comparatively easy. This suggests that reversed easy-task trajectories can serve as a scalable supervision signal for the hard task, reducing the cost of manual demonstration collection. However, reversed data can be noisy, and directly training on it may yield suboptimal policies. To enable largely automated acquisition and effective use of reversed data, we propose a teleoperation-cost effective framework for hard policy learning via temporal reversal of easy tasks, consisting of three key components: a closed-loop data collection pipeline that alternates between hard-task and easy-task policies to autonomously reset the environment and generate diverse trajectories; a hierarchical data refinement pipeline that temporally inverts easy-task rollouts and filters low-quality motion using kinematic priors and a critic-guided advantage filter; and an iterative policy learning method that trains the hard-task policy using both initial reversed easy-task demonstrations and the filtered reversed data in a continuous online learning loop. By combining automated collection, hierarchical refinement, and iterative learning, our method enables scalable, reliable training of complex, high-precision manipulation tasks. Across two simulated benchmarks and real-robot experiments, we demonstrate that our method improves hard-task success rates with higher data efficiency and more stable training compared to reversal-based and reinforcement-learning baselines, without requiring extensive hard-task teleoperation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。