arXiv:2605.27038cs.RO2026-05

让视觉语言模型更准预测车辆位置,减少自动驾驶事故

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

论文配图:TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving
图 1 · 摘自论文原文
  • 用任务引导的向量量化,把注意力从静态背景转向动态车辆
  • 在nuScenes数据集上降低碰撞率,闭路测试创安全新纪录
  • 适合做自动驾驶决策系统的研究人员和工程师

视觉语言模型(VLM)为自动驾驶规划提供了潜力,但如何将语义推理与精确的3D空间预测结合仍是挑战。现有方法或因将连续空间状态符号化而破坏几何结构,导致‘空间幻觉’;或因保留密集视觉信息而使背景纹理干扰模型,引发‘表示干扰’。为此,我们提出TPS-Drive,一种以任务为导向的表征净化框架,使VLM能在净化后的空间中思考。其核心是代理中心的分词器,通过冻结的3D检测头监督,将有限码本容量从冗余静态背景重分配至关键动态代理,有效隔离空间冗余。基于此净化后的空间词汇,TPS-Drive采用解耦推理流程,依次完成场景理解、未来预测与动作生成。通过渐进式三阶段训练与奖励驱动优化,显著优于纯模仿学习。大量实验验证:TPS-Drive实现精准的代理空间状态预测,在开环nuScenes评估中降低碰撞率,并在严苛的闭环NAVSIMv1和NAVSIMv2基准上创下新安全纪录。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial forecasting remains a critical challenge. Existing representation strategies generally follow two paths: text-aligned methods flatten continuous spatial states into symbols, which compromises geometric structure and induces "spatial hallucinations"; dense visual methods preserve spatial topology but overwhelm standard tokenizers with redundant background textures, leading to "representation interference". To address these limitations, we introduce TPS-Drive, a novel framework centered on Task-Guided Representation Purification that empowers VLMs to Think in Purified Space. At its core, an Agent-Centric Tokenizer utilizes a task-guided vector quantization mechanism supervised by a frozen 3D detection head, which explicitly reallocates limited codebook capacity from pervasive static backgrounds to critical dynamic agents and effectively isolates spatial redundancy. Leveraging this purified spatial vocabulary, TPS-Drive employs a decoupled reasoning pipeline that sequentially performs scene understanding, future forecasting, and action generation. The framework is optimized via a progressive three-stage training paradigm, culminating in reward-driven refinement that surpasses pure imitation learning. Extensive experiments validate our approach: TPS-Drive achieves accurate agent spatial state forecasting and reduces collision rates in open-loop nuScenes evaluations, while establishing new safety records on the rigorous closed-loop NAVSIMv1 and NAVSIMv2 benchmarks.

自动驾驶视觉语言模型空间预测表征净化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。