用残差流引导让机器人抓取更聪明,少试错就能学会复杂操作。
RFS: Reinforcement Learning with Residual Flow Steering for Dexterous Manipulation
- 通过残差动作+噪声调节,双路径探索提升适应效率
- 在仿真和真实机器人上均实现高效微调,仅需少量数据
- 适合想快速优化预训练抓取模型的研究者或工程师
模仿学习已成为机器人序列决策的有效方法,在高维灵巧操作任务中表现优异。近期行为克隆方法利用扩散模型、流匹配等生成模型表示多模态动作分布。但这类预训练策略泛化能力有限,部署时仍需额外微调以获得鲁棒性能。该过程需保留预训练的全局探索优势,同时快速修正局部执行误差。我们提出残差流引导(RFS)强化学习框架,用于高效适应预训练生成策略。RFS通过联合优化残差动作与潜在噪声分布,实现两种互补探索:通过残差修正进行局部精调,通过潜在空间调制实现全局探索。该设计在保持预训练策略表达结构的同时,实现了高效适应。我们在灵巧操作任务中验证了RFS有效性,在仿真与真实场景下均展现出高效的微调能力。
原文摘要 · Abstract (English)
Imitation learning has emerged as an effective approach for bootstrapping sequential decision-making in robotics, achieving strong performance even in high-dimensional dexterous manipulation tasks. Recent behavior cloning methods further leverage expressive generative models, such as diffusion models and flow matching, to represent multimodal action distributions. However, policies pretrained in this manner often exhibit limited generalization and require additional fine-tuning to achieve robust performance at deployment time. Such adaptation must preserve the global exploration benefits of pretraining while enabling rapid correction of local execution errors. We propose Residual Flow Steering(RFS), a data-efficient reinforcement learning framework for adapting pretrained generative policies. RFS steers a pretrained flow-matching policy by jointly optimizing a residual action and a latent noise distribution, enabling complementary forms of exploration: local refinement through residual corrections and global exploration through latent-space modulation. This design allows efficient adaptation while retaining the expressive structure of the pretrained policy. We demonstrate the effectiveness of RFS on dexterous manipulation tasks, showing efficient fine-tuning in both simulation and real-world settings when adapting pretrained base policies. Project website:https://weirdlabuw.github.io/rfs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。