arXiv:2607.29613cs.ROcs.CL2026-07

提出世界评论模型WCM,让机器人视觉-语言-动作系统更懂时间动态。

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

论文配图:WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
图 1 · 摘自论文原文
  • 用轻量LeJEPA架构联合预测未来状态与价值,显式建模时间结构
  • 在149个任务上超越现有方法,分布内外均表现优异
  • 兼容主流VLA模型,适合追求泛化能力的机器人研究者

视觉-语言-动作(VLA)模型的强化学习后训练在机器人操作中展现出巨大潜力。然而,现有基于评论家的方法多依赖单帧观测或单帧视觉模型隐状态,与机器人控制部分可观测的本质严重不符。简单引入历史信息会导致高维视觉空间中复杂度指数增长,且纯标量回报回归无法提供足够监督以学习跨时序动态。我们指出根本问题是状态近似:缺乏显式世界建模目标,导致评论家表示无法捕捉必要的时间结构。为此,提出世界评论模型(WCM),基于轻量LeJEPA架构,联合预测未来隐状态并估计价值,使评论家表示显式学习时间动态而非仅回归标量回报。WCM可无缝集成于在线与离线策略训练流程,兼容当前主流VLA骨干如Pi0、Pi0.5和OpenVLA-OFT。在四个基准上共149个任务的大量实验表明,WCM在分布内与分布外设置下均实现顶尖性能,尤其在泛化能力上提升显著。进一步在七项真实世界操作任务中使用OpenVLA-OFT与Pi0.5结合离线强化学习验证,确认其在多样化场景下的稳定部署能力。

原文摘要 · Abstract (English)

Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate observation history into the critic incurs exponential complexity with high-dimensional visual space, and still fails because pure scalar-return regression provides insufficient supervision for learning cross-temporal dynamics. We identify the root cause as a state approximation problem: without an explicit world modeling objective, the critic's representation cannot capture the temporal structure needed for accurate value estimation. To address this, we propose the World Critic Model (WCM), built on a lightweight LeJEPA architecture; WCM jointly predicts future latent state and estimates values, such that the critic's representation is explicitly trained to capture temporal dynamics rather than merely regress scalar returns. WCM integrates seamlessly into both on-policy and off-policy training pipelines and is compatible with state-of-the-art VLA backbones including Pi0, Pi0.5, and OpenVLA-OFT. Extensive experiments on 149 tasks across four benchmarks demonstrate that WCM consistently achieves state-of-the-art performance in both in-distribution and out-of-distribution settings, with particularly strong generalization gains. We further validate WCM on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5 with off-policy RL, confirming stable deployment across diverse settings.

强化学习机器人视觉语言动作时间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。