arXiv:2508.02219cs.ROcs.LG2025-08被引 37

用分块离线强化学习,高效微调视觉语言动作模型

CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning

  • 提出分块强化学习框架,支持动作分块的时序差分学习
  • 仅需30-60个示范样本,成功率提升57%,周期缩短22.3%
  • 适合需要少样本、强泛化能力的机器人控制场景

视觉语言动作(VLA)模型在实现真实世界机器人控制的通用策略方面展现出巨大潜力。尽管如此,使用强化学习(RL)对这类模型进行微调仍面临样本效率低、与动作分块不兼容及训练不稳定等挑战。为此,本文探索结合动作分块的离线强化学习来微调VLA模型。提出一种名为分块强化学习(Chunked RL)的新框架,将时序差分(TD)学习扩展至支持动作分块。基于此框架,提出CO-RFT算法,利用有限演示(30至60样本)微调VLA模型。首先通过全参数微调的模仿学习初始化骨干网络与策略;随后采用带动作分块的离线强化学习优化预训练策略。实验证明,该方法在真实环境中显著优于以往监督方法:成功率达57%提升,周期时间减少22.3%;且在未见过的位置上仍保持44.3%的成功率,展现鲁棒的位置泛化能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models demonstrate significant potential for developing generalized policies in real-world robotic control. This progress inspires researchers to explore fine-tuning these models with Reinforcement Learning (RL). However, fine-tuning VLA models with RL still faces challenges related to sample efficiency, compatibility with action chunking, and training stability. To address these challenges, we explore the fine-tuning of VLA models through offline reinforcement learning incorporating action chunking. In this work, we propose Chunked RL, a novel reinforcement learning framework specifically designed for VLA models. Within this framework, we extend temporal difference (TD) learning to incorporate action chunking, a prominent characteristic of VLA models. Building upon this framework, we propose CO-RFT, an algorithm aimed at fine-tuning VLA models using a limited set of demonstrations (30 to 60 samples). Specifically, we first conduct imitation learning (IL) with full parameter fine-tuning to initialize both the backbone and the policy. Subsequently, we implement offline RL with action chunking to optimize the pretrained policy. Our empirical results in real-world environments demonstrate that CO-RFT outperforms previous supervised methods, achieving a 57% improvement in success rate and a 22.3% reduction in cycle time. Moreover, our method exhibits robust positional generalization capabilities, attaining a success rate of 44.3% in previously unseen positions.

机器人控制强化学习少样本微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。