arXiv:2511.19528cs.ROcs.AI2025-11被引 2

用多策略强化学习生成多样化操作数据,提升视觉-语言-动作模型性能。

Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories

  • 基于信息论发现多种独立高成功率行为模式,突破单模式强化学习局限。
  • 在LIBERO数据集上生成轨迹多样性显著提升,覆盖更广状态-动作空间。
  • 预训练模型在未见任务上表现更优,适合构建可扩展的具身基础模型。

扩大视觉-语言-动作(VLA)模型预训练需要大量多样且高质量的操作轨迹。当前多数数据依赖人工遥控,成本高且难扩展。强化学习(RL)可通过自主探索学习有效技能,但标准RL训练易陷入单一执行模式,限制其大规模预训练应用。我们提出发现、学习与强化(DLR)框架,一种基于信息论的模式发现方法,可生成多种独特且高成功率的行为模式,用于VLA预训练。实验证明,DLR在LIBERO数据集上生成的轨迹库多样性显著提升:同一任务下,标准RL仅发现一种策略,而DLR能学习多个不同且成功率高的策略,从而覆盖更广泛的态-动空间。当适配未见过的下游任务套件时,基于多样化RL数据预训练的VLA模型性能优于等量标准RL数据训练的模型。此外,DLR表现出正向数据缩放特性,而单模式RL不具备此特性。这些结果表明,多模式强化学习是具身基础模型实用且可扩展的数据引擎。

原文摘要 · Abstract (English)

Scaling vision-language-action (VLA) model pre-training requires large volumes of diverse, high-quality manipulation trajectories. Most current data is obtained via human teleoperation, which is expensive and difficult to scale. Reinforcement learning (RL) methods learn useful skills through autonomous exploration, making them a viable approach for generating data. However, standard RL training collapses to a narrow execution pattern, limiting its utility for large-scale pre-training. We propose Discover, Lea rn and Reinforce (DLR), an information-theoretic pattern discovery framework that generates multiple distinct, high-success behavioral patterns for VLA pretraining. Empirically, DLR generates a markedly more diverse trajectory corpus on LIBERO. Specifically, it learns multiple distinct, high-success strategies for the same task where standard RL discovers only one, and hence it covers substantially broader regions of the state-action space. When adapted to unseen downstream task suites, VLA models pretrained on our diverse RL data surpass counterparts trained on equal-sized standard RL datasets. Moreover, DLR exhibits positive data-scaling behavior that single-pattern RL lacks. These results position multi-pattern RL as a practical, scalable data engine for embodied foundation models.

强化学习多策略具身智能预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。