arXiv:2605.12334cs.AI2026-05被引 1

用无任务世界模型实现零样本视觉语言动作适应,无需真实交互。

Reinforcing VLAs in Task-Agnostic World Models

论文配图:Reinforcing VLAs in Task-Agnostic World Models
图 1 · 摘自论文原文
  • 世界模型与奖励模型解耦,均基于无任务数据预训练。
  • 在仿真与真实场景中,新任务零样本适配性能提升显著。
  • 双噪声验证机制降低幻觉,适合快速部署新任务的系统。

通过在学习到的世界模型中进行强化学习,对后训练的视觉-语言-动作(VLA)模型进行微调,已成为在不依赖昂贵真实交互的前提下适应新任务的有效策略。然而,现有方法仍严重依赖特定任务数据来微调世界模型和奖励模型,从根本上限制了其对未见任务的可扩展性。为此,我们提出一种新范式 RAW-Dream(Reinforcing VLAs in task-Agnostic World Dreams),完全解耦世界模型学习与下游任务依赖。RAW-Dream 利用在多样化无任务行为上预训练的世界模型预测未来轨迹,并使用现成的视觉-语言模型(VLM)生成奖励。由于两个组件均为无任务导向,任何新任务均可在此零样本想象空间内完成微调。此外,为缓解世界模型幻觉问题,引入双噪声验证机制以过滤不可靠轨迹。在仿真与真实世界设置中的大量实验表明,性能持续提升,证明通用物理先验可有效替代高成本的任务依赖数据,为 VLA 适应提供了高度可扩展的路径。

原文摘要 · Abstract (English)

Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to adapt to new tasks without costly real-world interactions. However, while using imagined trajectories reduces the sample complexity of policy training, existing methods still heavily rely on task-specific data to fine-tune both the world and reward models, fundamentally limiting their scalability to unseen tasks. To overcome this, we argue that world and reward models should capture transferable physical priors that enable zero-shot inference. We propose RAW-Dream (Reinforcing VLAs in task-Agnostic World Dreams), a new paradigm that completely disentangles world model learning from downstream task dependencies. RAW-Dream utilizes a world model pre-trained on diverse task-free behaviors for predicting future rollouts, and an off-the-shelf Vision-Language Model (VLM) for reward generation. Because both components are task-agnostic, VLAs can be readily finetuned for any new task entirely within this zero-shot imagination. Furthermore, to mitigate world model hallucinations, we introduce a dual-noise verification mechanism to filter out unreliable rollouts. Extensive experiments across simulation and real-world settings demonstrate consistent performance gains, proving that generalized physical priors can effectively substitute for costly task-dependent data, offering a highly scalable roadmap for VLA adaptation.

视觉语言动作世界模型零样本强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。