arXiv:2607.20988cs.CVcs.AI2026-07

融合像素与隐空间建模,提升自动驾驶模型在噪声环境下的鲁棒性。

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

论文配图:HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving
图 1 · 摘自论文原文
  • 双阶段训练:先重建图像帧,再预测视频隐变量。
  • 在NAV SIM v1/v2上性能超越纯像素或纯隐空间模型。
  • 首次系统评估世界模型噪声鲁棒性,适合自动驾驶研究者。

视觉-语言-动作(VLA)模型结合世界建模为端到端自动驾驶提供了新范式。虽然像素级未来预测支持精细时空推理,但在嘈杂驾驶场景中鲁棒性差;而基于隐空间的世界模型虽更稳定,却因缺乏像素级对齐导致可解释性下降和表征退化。为此,我们提出HyWorldVLA,一种统一像素级监督与隐表示学习的混合世界建模框架。预训练阶段,模型同时预测由预训练视频变分自编码器(video VAE)编码的视频隐变量,并重建视频帧以提供精确像素级对齐。在后续共微调阶段,模型仅预测隐变量,并输入动作专家生成轨迹。在NAV SIM v1和v2基准上的大量实验表明,HyWorldVLA显著优于基于像素和基于隐空间的世界模型基线。特别地,我们首次对自动驾驶中的世界模型噪声鲁棒性进行了全面的定性和定量分析,建立了评估未来架构的新基准。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning. In the pre-training stage, HyWorldVLA predicts video latents encoded by a pre-trained video VAE, while simultaneously reconstructing video frames to provide precise pixel-level grounding. During the subsequent co-fine-tuning phase, the model exclusively predicts latent features, which are fed into an action expert to generate trajectories. Extensive experiments on NAVSIM v1 and v2 benchmarks demonstrate that HyWorldVLA significantly outperforms both pixel-based and latent-based world model baselines. Notably, we present the first comprehensive qualitative and quantitative analysis of world model noise robustness in autonomous driving, establishing a new benchmark for evaluating future architectures.

自动驾驶世界建模VLA视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。