arXiv:2510.06710cs.RO2025-10中稿 · RSS 2026被引 14

统一高效训练视觉-语言-动作模型的强化学习框架,显著提升训练速度与性能。

RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models

  • 构建统一接口整合多种模型、算法和仿真器,支持可复现研究。
  • 在ManiSkill上实现1.61至1.88倍训练加速,多任务成功率超97%。
  • 适合从事具身智能、强化学习与多模态模型研究的团队使用。

近期研究表明,通过交互式强化学习(RL)可提升视觉-语言-动作(VLA)模型的任务表现。然而,现有工作仍分散,缺乏统一平台用于公平比较不同架构与算法,也缺少可扩展的高效系统设计。为此,我们提出RLinf-VLA——一个面向VLA模型可扩展强化学习训练的统一高效框架。该框架通过统一接口标准化集成多样VLA架构、强化学习算法及异构仿真环境,支持可扩展性与可复现性。为提升效率,框架采用灵活资源分配架构,针对渲染、推理与训练流程进行优化。特别地,引入混合细粒度流水线分配策略,在ManiSkill上实现1.61×至1.88×的训练速度提升。基于此框架,训练模型在多个具身基准测试中表现优异:在130个LIBERO任务上达成98.11%成功率,在25个ManiSkill任务上达97.66%,在6个RoboTwin任务上平均成功率84.63%。此外,框架提炼出一套有效的基于RL的VLA训练实践。我们期望RLinf-VLA成为具身智能领域高效、统一、可复现研究的基础框架。

原文摘要 · Abstract (English)

Recent studies have demonstrated the potential of reinforcement learning (RL) to improve the task performance of vision-language-action (VLA) models through interaction. However, current efforts remain fragmented, lacking a unified platform for fair comparison across architectures and algorithms, as well as an efficient system design for scalable training. Therefore, we present RLinf-VLA, a unified and efficient framework for scalable RL training of VLA models. RLinf-VLA standardizes the integration of diverse VLA architectures, RL algorithms, and heterogeneous simulators through a unified interface, enabling extensibility and reproducibility. To improve efficiency, the framework adopts a flexible resource allocation architecture for rendering, inference, and training in RL pipelines. In particular, RLinf-VLA introduces a hybrid fine-grained pipeline allocation strategy that achieves a 1.61$\times$-1.88$\times$ training speedup on ManiSkill. Using this framework, RL-trained models achieve strong performance across embodied benchmarks, including 98.11% success on 130 LIBERO tasks, 97.66% success on 25 ManiSkill tasks, and 84.63% average success across 6 RoboTwin tasks. In addition, RLinf-VLA distills a set of effective practices for RL-based VLA training. We envision RLinf-VLA as a foundational framework for efficient, unified, and reproducible research in embodied intelligence.

强化学习具身智能多模态训练框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。