让机器人学会又快又稳地执行复杂操作,通过强化学习自动优化动作节奏。
SpeedAug: Policy Acceleration via Tempo-Enriched Policy and RL Fine-Tuning
- 用加速演示数据训练带节奏的初始策略,覆盖多种执行速度。
- 基于强化学习微调策略,高效优化动作轨迹和执行节奏。
- 实测仅需16分钟在线交互,任务吞吐量提升1.8倍,成功率不降。
复杂真实世界操作任务的机器人策略学习近年来进展迅速,主要得益于通过人工操作收集示范数据。然而,基于此类示范训练的策略往往执行速度远低于机器人的物理能力,因为示范数据在实际约束下更倾向于保守、确保成功的行为,而非追求速度。现有策略加速方法依赖数据预处理或启发式规则确定执行节奏,而非学习任务最优的执行速度。本文提出 SpeedAug 框架,通过强化学习使策略自主学习任务最优执行节奏。SpeedAug 首先从速度增强的示范数据中学习一个包含多样执行节奏的先验策略;在此基础上,利用强化学习微调引导探索,高效优化动作轨迹并优化执行节奏。在多个机器人操作基准测试中,SpeedAug 显著提升了策略加速的样本效率,同时保持高成功率,实现快速且稳定的任务执行。应用于真实世界操作任务时,仅需16分钟在线交互,任务吞吐量提升1.8倍,且成功率不受影响。
原文摘要 · Abstract (English)
Robotic policy learning for complex real-world manipulation tasks has seen rapid recent progress, enabled in large part by the ability to collect demonstrations through human operation. However, policies trained from such demonstrations often execute tasks far more slowly than the robot's physical capabilities, as demonstration data is collected under practical constraints that favor conservative, success-oriented trajectories over execution speed. Existing policy acceleration methods determine execution tempo through data preprocessing or heuristic rules, rather than learning execution speed optimized for the task. In this paper, we propose SpeedAug, a policy acceleration framework that enables policies to learn task-optimal execution tempo via reinforcement learning (RL). SpeedAug first learns a tempo-enriched prior policy from speed-augmented demonstrations that captures diverse execution tempos. Building on this tempo-enriched prior, RL fine-tuning guides exploration to refine action trajectories and optimize execution tempo efficiently. Experiments on robotic manipulation benchmarks demonstrate that SpeedAug substantially improves the sample efficiency of policy acceleration while maintaining high success rates, achieving fast and stable task execution. Applied to a real-world manipulation task, SpeedAug improves task throughput by 1.8x using only 16 minutes of online interactions without compromising the success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。