让自动驾驶模型在隐空间中进行物理可信的时空推理
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
- 在隐式时空空间中进行连续推理,融合几何约束与动态预判
- 在NAVSIM v1和v2上分别达到91.3和87.1的领先指标
- 适合需要高安全性和复杂时空决策的自动驾驶系统
尽管视觉-语言-行动(VLA)模型通过统一感知与规划推动了自动驾驶的发展,但其依赖显式文本思维链(CoT)导致语义与感知脱节、符号与感知冲突。近期转向隐式推理虽绕过瓶颈,但缺乏显式中间约束,常表现为无物理依据的表征。为此,我们提出潜空间时序视觉-语言-行动框架(LaST-VLA),将推理范式从离散符号处理转变为具有物理基础的潜空间时序思维链。通过双特征对齐机制,从3D基础模型中提取几何约束,从世界模型中注入动态预见能力,直接融入潜空间。结合渐进式监督微调策略,从特征对齐逐步过渡到轨迹生成,并通过组相对策略优化(GRPO)强化训练,确保安全性与规则合规性。该方法在NAVSIM v1(91.3 PDMS)和NAVSIM v2(87.1 EPDMS)上刷新纪录,同时在SURDS与NuDynamics基准上展现出卓越的时空推理能力。
原文摘要 · Abstract (English)
While Vision-Language-Action (VLA) models have revolutionized autonomous driving by unifying perception and planning, their reliance on explicit textual Chain-of-Thought (CoT) leads to semantic-perceptual decoupling and perceptual-symbolic conflicts. Recent shifts toward latent reasoning attempt to bypass these bottlenecks by thinking in continuous hidden space. However, without explicit intermediate constraints, standard latent CoT often operates as a physics-agnostic representation. To address this, we propose the Latent Spatio-Temporal VLA (LaST-VLA), a framework shifting the reasoning paradigm from discrete symbolic processing into a physically grounded Latent Spatio-Temporal CoT. By implementing a dual-feature alignment mechanism, we distill geometric constraints from 3D foundation models and dynamic foresight from world models directly into the latent space. Coupled with a progressive SFT training strategy that transitions from feature alignment to trajectory generation, and refined via Reinforcement Learning with Group Relative Policy Optimization (GRPO) to ensure safety and rule compliance. \method~setting a new record on NAVSIM v1 (91.3 PDMS) and NAVSIM v2 (87.1 EPDMS), while excelling in spatial-temporal reasoning on SURDS and NuDynamics benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。