让自动驾驶模型看清全局路况,精准规划路径。
DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

- 用俯视图增强语言模型的空间感知能力
- 实拍与渲染图像对齐,提升视觉注意力精度
- 自评优化轨迹选择,适合高阶自动驾驶研究
视觉-语言-动作(VLA)驾驶模型将预训练的视觉-语言模型转化为驾驶策略,可利用世界知识并遵循语言指令。然而,现有VLA模型缺乏面向驾驶的空间智能:其策略主要基于视角图像令牌和语言先验,而精确运动规划需要度量几何、俯视场景结构以及对安全关键感知线索的关注。这一局限导致模型易受专家演示中视觉几何建模弱和感知覆盖不足的影响。本文提出DriveStack-VLA,基于大型VLM主干网络构建。为强化VLA驾驶的空间定位,设计双视觉建模组件:通过DeepStack式连接将鸟瞰图注入大语言模型解码器,并提出渲染教师对齐方法,使真实图像与栅格化图像的感知焦点对齐。此外,为弥合多模态轨迹选择差距,引入基于头部的自批判模块,对采样轨迹排序并有条件地优化最优轨迹。DriveStack-VLA在NAVSIMv1上达到91.6 PDMS,NAVSIMv2上达91.0 EPDMS(启用人类惩罚过滤器),在闭环Bench2Drive上取得79.49分驾驶评分与56.36%成功率。更多可视化见项目页:https://anonymous.4open.science/w/drivestack-vla/
原文摘要 · Abstract (English)
Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow language guidances. However, existing VLA driving models still lack driving-oriented spatial intelligence: their policies are mainly grounded on perspective image tokens and language priors, while precise motion planning requires metric geometry, top-down scene structure, and attention to safety-critical perceptual cues. This limitation makes current models vulnerable to weak visual geometry modeling and perceptual coverage in expert demonstrations. In this paper, we present DriveStack-VLA, a framework built upon a large VLM backbone. To strengthen the spatial grounding of VLA driving, we develop dual visual modeling components. We inject a Bird-Eye-View representation into the Large Language Model decoder through a DeepStack-style connection, and propose Render-Teacher Alignment to align the perceptual focus of real images with that of rasterized images. Furthermore, to bridge the gap in multimodal trajectory selection, we introduce a head-based self-critique module that ranks sampled trajectories and conditionally refines the best one. DriveStack-VLA achieves 91.6 PDMS on NAVSIMv1, 91.0 EPDMS on NAVSIMv2 (with the human penalty filter enabled), and a driving score of 79.49 with a success rate of 56.36\% on the closed-loop Bench2Drive. More visualizations are available on our project page: https://anonymous.4open.science/w/drivestack-vla/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。