arXiv:2603.22078cs.RO2026-03被引 16

对比世界模型与视觉语言模型在机器人任务中的泛化能力。

Do World Action Models Generalize Better than VLAs? A Robustness Study

  • 用视频预训练的世界动作模型直接预测环境动态并生成动作。
  • 世界动作模型在扰动下表现更稳定,最高成功率82.2%。
  • 适合关注机器人鲁棒性与少样本泛化的研究者。

真实世界中的机器人动作规划面临挑战,需理解环境状态并预测其响应动作的演化。视觉-语言-动作(VLA)模型通过复用大规模视觉-语言模型并结合动作专家,在多种机器人任务中取得显著进展,但其性能受限于训练数据范围,对未见场景泛化能力弱,且易受上下文扰动影响。近期,世界模型作为替代方案被重新关注,称为世界动作模型(WAM),基于大规模视频数据预训练的世界模型可预测未来状态,经微调后能解码为机器人动作。有观点认为其显式动态预测能力与从网络规模视频中学习的空间时间先验,使其比VLA更具泛化优势。本文对主流SOTA VLA策略与新发布的WAM进行对比研究,评估其在LIBERO-Plus与RoboTwin 2.0-Plus基准上的表现,涵盖多种视觉与语言扰动。结果表明,WAMs展现出强鲁棒性:LingBot-VA在RoboTwin 2.0-Plus上达74.2%成功率,Cosmos-Policy在LIBERO-Plus上达82.2%。尽管某些VLA如$π_{0.5}$在特定任务上表现相当,但通常需大量多样机器人数据和复杂训练目标。相关评估代码已公开于https://robot-robustness.github.io/RoboTwin2.0-Plus/。

原文摘要 · Abstract (English)

Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response to actions. Vision-language-action (VLA), which repurpose large-scale vision-language models for robot action generation using action experts, have achieved notable success across a variety of robotic tasks. Nevertheless, their performance remains constrained by the scope of their training data, exhibiting limited generalization to unseen scenarios and vulnerability to diverse contextual perturbations. More recently, world models have been revisited as an alternative to VLAs. These models, referred to as world action models (WAMs), are built upon world models that are trained on large corpora of video data to predict future states. With minor adaptations, their latent representation can be decoded into robot actions. It has been suggested that their explicit dynamic prediction capacity, combined with spatiotemporal priors acquired from web-scale video pretraining, enables WAMs to generalize more effectively than VLAs. In this paper, we conduct a comparative study of prominent state-of-the-art VLA policies and recently released WAMs. We evaluate their performance on the LIBERO-Plus and RoboTwin 2.0-Plus benchmarks under various visual and language perturbations. Our results show that WAMs achieve strong robustness, with LingBot-VA reaching 74.2% success rate on RoboTwin 2.0-Plus and Cosmos-Policy achieving 82.2% on LIBERO-Plus. While VLAs such as $π_{0.5}$ can achieve comparable robustness on certain tasks, they typically require extensive training with diverse robotic datasets and varied learning objectives. The evaluation code for the RoboTwin2.0-Plus benchmark is available at: https://robot-robustness.github.io/RoboTwin2.0-Plus/.

机器人世界模型泛化能力鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。