为自动驾驶设计可同时支持建模与规划的离散视觉令牌化方法
Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning

- 联合监督下学习,对齐冻结DINO特征空间并保留外观
- 引入相邻帧深度与相对位姿监督,提升几何一致性
- 适合需要高效建模与决策的自动驾驶系统
离散视觉令牌应为基于令牌的世界建模与规划提供紧凑表示。然而,大多数令牌化方法源自图像生成,主要优化像素重建,可能在易生成与驱动决策有用性之间存在差距。本文提出一种表征引导且几何增强的令牌化方法,在联合监督下学习离散令牌。通过特征解码对齐冻结DINO特征空间,同时利用感知与对抗损失进行RGB重建以保留外观。训练中加入相邻帧深度与相对位姿监督,并通过多码本量化稳定联合目标。在NAVSIM上评估,相同学习令牌在轻量级规划读出与GPT风格的下一令牌世界模型中均表现更优:重建保真度更高,表征一致性更强,在固定解码器下规划性能具有竞争力,匹配设置下生成质量更佳。
原文摘要 · Abstract (English)
Discrete visual tokens should provide a compact representation for both token-based world modeling and planning in autonomous driving. However, most tokenizers are inherited from image generation and are optimized mainly for pixel reconstruction, which may leave a gap between what is easy to generate and what is useful to decode for driving decisions. We present a representation-guided and geometry-enhanced tokenizer that learns discrete tokens under joint supervision. The tokenizer aligns its discrete bottleneck with a frozen DINO feature space through feature decoding, while preserving appearance via RGB reconstruction with perceptual and adversarial losses. To inject geometric state-related cues, we add adjacent-frame depth and relative-pose supervision during training and stabilize joint objectives with multi-codebook quantization. We evaluate the same learned tokens with a lightweight planning readout and a GPT-style next-token world model. Experiments on NAVSIM show improved reconstruction fidelity and representation consistency, competitive planning performance under a fixed decoder, and better generative quality under matched settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。