用稀疏语义场实现开放词汇的自动驾驶,兼顾效率与泛化能力。
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving

- 基于视觉语言模型生成无类别依赖的语义提示,构建稀疏能量场。
- 在nuScenes和CODA数据集上均实现高精度轨迹规划与碰撞规避。
- 适合研究开放世界自动驾驶与可解释性控制的开发者。
将端到端自动驾驶扩展至复杂开放环境,需具备对异常场景的泛化感知能力与生成符合车辆运动学约束轨迹的规划能力。现有范式在表征效率与泛化能力间存在显著矛盾:密集模型(如占据网络)几何鲁棒性强,但计算开销大,且难以进行高层语义推理;而稀疏查询式规划器虽高效,却依赖封闭集定义,易受分布外(OOD)事件影响。尽管近期视觉-语言-动作(VLA)模型提供开放词汇推理能力,其自回归离散令牌生成机制与车辆动力学所需的连续高频控制存在根本冲突。为此,我们提出Lagrange,一种基于掩码潜空间场(MLF)的开放词汇、计算稀疏驾驶框架。不依赖密集体素重建或封闭集查询机制,Lagrange利用视觉语言模型(VLM)将无类别依赖的对象提案编码为连续语义视觉令牌。引入意图驱动的掩码交叉注意力模块,动态过滤无关实体,将注意力后的令牌解码为定义在空间坐标上的隐式连续能量场。通过将决策建模为跨越该能量场的拉格朗日动作最小化问题,强制满足车辆运动学约束并实现避障。在标准(nuScenes)与长尾(CODA)基准上的大量离线评估表明,Lagrange建立了稳健、可解释且运动学可行的开放世界自主驾驶框架。
原文摘要 · Abstract (English)
Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories. Existing paradigms face a distinct dichotomy between representational efficiency and generalization capacity. Dense models (e.g., occupancy networks), while geometrically robust, incur critical computational bottlenecks and struggle with high-level semantic reasoning. Conversely, sparse, query-based planners are efficient but reliant on closed-set definitions, rendering them vulnerable to out-of-distribution (OOD) events. Although recent Vision-Language-Action (VLA) models offer open-vocabulary reasoning, their autoregressive, discrete token generation fundamentally conflicts with the continuous, high-frequency control requirements of vehicle dynamics. To address this, we propose Lagrange, an open-vocabulary, computationally sparse driving framework based on Masked Latent Fields (MLF). Rather than relying on dense volumetric reconstructions or closed-set query mechanisms, Lagrange exploits Vision-Language Models (VLMs) to encode class-agnostic object proposals into continuous semantic visual tokens. We introduce an intent-driven masked cross-attention module that temporally filters irrelevant entities, decoding the attended tokens into an implicit continuous energy field defined over spatial coordinates. By framing decision-making as a Lagrangian action minimization problem spanning this energy field, we enforce strict compliance with vehicle kinematics while executing collision avoidance. Extensive offline evaluations on both standard (nuScenes) and long-tail (CODA) benchmarks demonstrate that Lagrange establishes a promising framework for robust, interpretable, and kinematically feasible open-world autonomy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。