构建分层时空视觉语言动作模型,提升自动驾驶轨迹生成的准确性与鲁棒性。
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
- 分层时空架构融合几何感知与驾驶指令,增强三维空间和时间推理能力。
- 在NAVSIM v2基准上实现88.6的EPDMS(Navtest)与50.9的EPDMS(Navhard)。
- 动态隐变量正则化确保语言指令空间精准对齐,适合高安全要求自动驾驶场景。
视觉-语言-动作(VLA)模型通过多模态理解为自动驾驶提供了潜力,但在安全关键场景中受限于精确数值推理不足、三维空间感知弱以及对上下文敏感等问题。为此,我们提出HiST-VLA,一种专为可靠轨迹生成设计的分层时空视觉语言动作模型。该框架通过融合几何感知、细粒度驾驶指令与状态历史提示,增强三维空间与时间推理能力。为保证计算效率,将动态令牌稀疏化集成至VLA架构中,通过融合冗余令牌而非过滤,有效降低冗余且不牺牲性能。此外,采用基于分层Transformer的规划器,逐步将粗略的VLA航点细化为精细轨迹。关键在于,规划器利用动态潜在正则化融入语言指令,确保严格的空间定位与时间连贯性。在NAVSIM v2基准上的广泛评估显示,该模型在Navtest任务上达到88.6的EPDMS,在伪闭环的Navhard任务上达到50.9的EPDMS,表现优于现有方法。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models offer promising capabilities for autonomous driving through multimodal understanding. However, their utilization in safety-critical scenarios is constrained by inherent limitations, including imprecise numerical reasoning, weak 3D spatial awareness, and high sensitivity to context. To address these challenges, we propose HiST-VLA, a novel Hierarchical Spatio-Temporal VLA model designed for reliable trajectory generation. Our framework enhances 3D spatial and temporal reasoning by integrating geometric awareness with fine-grained driving commands and state history prompting. To ensure computational efficiency, we integrate dynamic token sparsification into the VLA architecture. This approach fuses redundant tokens rather than filtering them, effectively reducing redundancy without sacrificing model performance. Furthermore, we employ a hierarchical transformer-based planner to progressively refine coarse VLA waypoints into fine-grained trajectories. Crucially, the planner utilizes dynamic latent regularization to incorporate language commands, ensuring strict spatial grounding and temporal coherence. Extensive evaluation on the NAVSIM v2 benchmark demonstrates state-of-the-art performance on Navtest, achieving an EPDMS of 88.6, and EPDMS of 50.9 on pseudo closed-loop Navhard benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。