让自动驾驶同时提升决策与语言理解能力
ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
- 通过视觉与语言跨模态对齐,统一感知、预测与规划链路
- 在四个基准上同时刷新驾驶与语言推理的最优表现
- 训练时对齐,推理零开销,适合追求端到端智能的开发者
近期研究尝试将大语言模型(LLMs)融入端到端自动驾驶系统以增强泛化与可解释性,但多数方法仅聚焦于驾驶性能或视觉-语言推理,难以兼顾二者。本文提出ALN-P3,一种统一的联合蒸馏框架,实现快速视觉驱动系统与慢速语言推理模块间的跨模态对齐。ALN-P3引入三种新对齐机制:感知对齐(P1A)、预测对齐(P2A)和规划对齐(P3A),显式对齐感知、预测与规划全链条中的视觉标记与对应语言输出。所有对齐模块仅在训练阶段使用,推理时无额外开销。在nuScenes、Nu-X、TOD3Cap和nuScenes QA四个挑战性基准上的大量实验表明,ALN-P3显著提升驾驶决策与语言推理能力,达到当前最优水平。
原文摘要 · Abstract (English)
Recent advances have explored integrating large language models (LLMs) into end-to-end autonomous driving systems to enhance generalization and interpretability. However, most existing approaches are limited to either driving performance or vision-language reasoning, making it difficult to achieve both simultaneously. In this paper, we propose ALN-P3, a unified co-distillation framework that introduces cross-modal alignment between "fast" vision-based autonomous driving systems and "slow" language-driven reasoning modules. ALN-P3 incorporates three novel alignment mechanisms: Perception Alignment (P1A), Prediction Alignment (P2A), and Planning Alignment (P3A), which explicitly align visual tokens with corresponding linguistic outputs across the full perception, prediction, and planning stack. All alignment modules are applied only during training and incur no additional costs during inference. Extensive experiments on four challenging benchmarks-nuScenes, Nu-X, TOD3Cap, and nuScenes QA-demonstrate that ALN-P3 significantly improves both driving decisions and language reasoning, achieving state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。