用高层语义表征增强视觉语言动作模型,提升自动驾驶决策能力
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

- 在潜在空间中学习未来场景表征,避免像素级重建干扰
- 单次前向传播联合预测场景与轨迹,效率更高
- 两阶段轨迹解码利用未来表征优化路径生成,适合复杂路况
视觉-语言-动作(VLA)模型为端到端自动驾驶提供了新范式。现有方法多依赖稀疏动作监督,未能充分发挥其场景理解与推理能力。近期尝试通过世界建模引入密集视觉监督,但常过度关注像素级图像重建,忽视语义场景表征学习。本文提出LVDrive,一种基于潜在视觉表征的VLA框架。该框架将未来场景预测任务引入VLA范式,未来表示在高层潜在空间中学习,由预训练视觉主干提供辅助监督。不同于低效的自回归生成,我们以统一嵌入空间联合建模未来场景与运动预测,单次前向传播完成前瞻性推理。进一步设计两阶段轨迹解码策略,显式利用学习到的潜在未来表示优化轨迹生成。在Bench2Drive挑战性基准上的大量实验表明,LVDrive在闭环驾驶性能上显著优于动作监督方法与基于图像重建的世界模型方法。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and reasoning capabilities. Recent attempts to incorporate dense visual supervision via world modeling often overemphasize pixel-level image reconstruction, neglecting semantically meaningful scene representation learning. In this work, we propose LVDrive, a Latent Visual representation enhanced VLA framework for autonomous driving. LVDrive introduces a future scene prediction task into the VLA paradigm, where future representations are learned entirely in a high-level latent space under auxiliary supervision from a pretrained vision backbone. Departing from inefficient autoregressive generation, we jointly model future scene and motion prediction within a unified embedding space, processed in a single forward pass to conduct the future-aware reasoning. We further design a two-stage trajectory decoding strategy that explicitly leverages the learned latent future representations to refine trajectory generation. Extensive experiments on the challenging Bench2Drive benchmark demonstrate that LVDrive achieves significant improvements in closed-loop driving performance, outperforming both action supervised methods and image-reconstruction-based world model approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。