通过重建图像块特征提升驾驶场景令牌的表征能力
NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving

- 用自蒸馏掩码重建任务约束场景令牌的视觉信息
- 在Waymo上达8.0461 RFS,NavSim上94.1 PDMS
- 无需额外感知头,适合部署于端到端自动驾驶系统
近期无感知端到端自动驾驶方法通过将密集图像块令牌压缩为紧凑的场景令牌,直接用于轨迹生成与评分。尽管这些场景令牌构成规划器的紧凑视觉瓶颈,但仅受规划目标监督,对编码的视觉信息约束有限。为此,我们提出神经令牌重建(NTR),一种直接约束无感知驾驶中场景令牌瓶颈的表征学习框架。NTR引入自蒸馏掩码潜在重建目标,仅使用紧凑场景令牌作为重建记忆来恢复被掩码的块级潜在特征。这使得重建梯度仅通过场景令牌瓶颈传递,促使场景令牌保留更丰富且冗余更低的视觉表示。我们进一步引入基于基础模型标注的语义先验,作为弱语义接口,使重建目标偏向驾驶相关结构,而不引入显式感知头。所有辅助重建组件在推理时移除,部署的规划器保持不变。NTR在三个公开自动驾驶基准上达到最先进性能,包括Waymo E2E上的8.0461 RFS,以及NavSim1&2上的94.1 PDMS / 90.9 EPDMS。学习到的场景令牌表现出更低的成对冗余和更高的有效秩,表明有效的瓶颈监督同时提升了紧凑视觉表征学习与规划性能。
原文摘要 · Abstract (English)
Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compact scene tokens for downstream trajectory generation and scoring. While these scene tokens form a compact visual bottleneck for the planner, they receive supervision solely from the planning objective, providing limited constraints on the encoded visual information. To address this limitation, we introduce Neural Token Reconstruction (NTR), a representation learning framework to directly constrain the compact scene-token bottleneck in perception-free driving. NTR introduces a self-distillation masked latent reconstruction objective that reconstructs masked patch-level latent features using only compact scene tokens as reconstruction memory. This forces reconstruction gradients to pass exclusively through the scene-token bottleneck, encouraging scene tokens to preserve richer and less redundant visual representations for planning. We further introduce semantic priors derived from foundation-model annotations as a weak semantic interface biasing reconstruction targets toward driving-related structures without introducing explicit perception heads. All auxiliary reconstruction components are removed at inference time, leaving the deployed planner unchanged. NTR achieves state-of-the-art performance on three public autonomous driving benchmarks, including 8.0461 RFS on Waymo E2E and 94.1 PDMS / 90.9 EPDMS on NavSim1&2. The learned scene tokens exhibit lower pairwise redundancy and higher effective rank, indicating that effective bottleneck supervision improves both compact visual representation learning and planning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。