arXiv:2607.19399cs.LGcs.AI2026-07

用离线监督提升视觉语言动作模型的强化学习效率与泛化能力

Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models

论文配图:Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
图 1 · 摘自论文原文
  • 结合离线数据或预训练参考策略指导强化学习
  • 仅需一半训练预算即可接近标准强化学习性能
  • 兼顾高效训练与强泛化能力,适合大规模智能体训练

在线强化学习(RL)在多数评估指标上优于离线方法,尤其在分布外(OOD)场景表现更佳。最近研究提出一个专注于OOD的基准测试,发现经强化学习训练的视觉-语言-动作(VLA)策略在OOD和分布内(IND)表现均优于仅使用监督微调(SFT)的方法。本文探索混合离线-在线训练能否融合两者优势:通过离线数据或离线训练的参考策略对强化学习进行正则化。在该OOD基准上的实验表明,单纯离线训练无法实现良好OOD性能,但将离线监督引入强化学习后,既保持了强泛化能力,又显著提升了训练效率。具体而言,引导型方法在约一半训练预算下达到接近标准强化学习的性能,实现了效率与泛化的双赢。项目页面:https://alstar8.github.io/offline-supervision-vla-rl

原文摘要 · Abstract (English)

It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger out-of-distribution (OOD) behavior, where models trained only with imitation learning approaches often struggle. A recent study introduced an OOD-focused benchmark and reported that RL-trained vision-language-action (VLA) policies achieve noticeably better OOD performance and slightly better in-distribution (IND) performance than their counterparts trained with supervised fine-tuning (SFT). In this work, we investigate whether hybrid offline-online training can combine the advantages of both approaches. Specifically, we study RL methods regularized by offline supervision via either offline data or an offline-trained reference policy. We evaluate these approaches on the OOD benchmark and compare them with both offline-only training and standard RL. Our results show that although offline training achieves limited OOD performance by itself, incorporating offline supervision into RL preserves strong OOD capability while substantially improving training efficiency. In particular, the guided methods reach performance close to that of standard RL while requiring roughly half of the training budget. Rather than producing a trade-off between speed and OOD performance, the hybrid approach retains strong OOD capability while achieving this efficiency gain. Project page: https://alstar8.github.io/offline-supervision-vla-rl

强化学习视觉语言动作离线监督高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。