arXiv:2607.03182cs.ROcs.AI2026-07

用锚点连接语义决策与连续轨迹,提升自动驾驶规划效率与准确性。

AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

论文配图:AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning
图 1 · 摘自论文原文
  • 以锚点表示驾驶行为模式,替代逐点生成轨迹
  • 在真实场景中实现77.28%成功率和89.92分驾驶得分
  • 适合需要高语义理解与高效推理的自动驾驶系统

自动驾驶规划需将导航意图、交通规则、动态交互和语言指令转化为可执行的连续轨迹。视觉-语言-动作(VLA)模型被引入以增强长尾泛化、常识推理、高层语义理解和可解释性。然而,现有VLA规划器多采用基于规划头的轨迹预测或全序列自回归生成,前者对连续轨迹约束弱,后者依赖低信息密度的坐标令牌序列,导致语义-动作对齐困难、离散化误差大且推理效率低。为此,我们提出层级式决策锚定框架AnchorVLA,通过轨迹模式锚点作为高层VLA推理与连续轨迹执行之间的显式接口。具体而言,决策锚定表示法以锚点编码完整局部运动模式,而非单个坐标点;决策锚定残差流在选定锚点定义的残差空间中生成精细连续轨迹,捕捉高层决策后的多模态执行优化。相比自回归生成路径序列,AnchorVLA在保留大语言模型决策能力的同时,提升了推理效率、语义-动作对齐度与连续生成灵活性。在Bench2Drive闭环基准测试中,AnchorVLA达到77.28%的最优成功率和89.92的竞争力驾驶得分。

原文摘要 · Abstract (English)

Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable continuous trajectories. Vision-Language-Action models have been introduced into driving planning to improve long-tail generalization, commonsense reasoning, high-level semantic understanding, and explainability. However, existing VLA planners mainly follow planning-head-based trajectory prediction or full-trajectory autoregressive generation. The former only weakly constrains continuous trajectory generation with VLA reasoning, while the latter relies on long sequences of low-information-density coordinate tokens, making semantic-action alignment difficult and leading to discretization errors and inefficient inference. To address these limitations, we propose AnchorVLA, a hierarchical decision-anchored VLA planning framework that uses trajectory-pattern anchors as an explicit interface between high-level VLA reasoning and continuous trajectory execution. Specifically, Decision-as-Anchor Representation represents behavior-level driving decisions with anchor tokens, each encoding an entire local motion pattern rather than a single coordinate point. Decision-Anchored Residual Flow then generates fine-grained continuous trajectories in the selected anchor-defined residual space, capturing multi-modal execution refinements after high-level decision making. By reasoning over compact and semantically meaningful anchors instead of autoregressively generating waypoint sequences, AnchorVLA preserves LLM-based decision making while improving inference efficiency, semantic-action alignment, and continuous generation flexibility. Experiments on the Bench2Drive closed-loop benchmark show that AnchorVLA achieves a state-of-the-art Success Rate of 77.28 and a competitive Driving Score of 89.92.

自动驾驶视觉语言动作轨迹生成决策锚定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。