让自动驾驶模型先想后行,动态修正决策更可靠
Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving

- 用文本和鸟瞰图双重预测反思未来场景变化
- 在NAVSIM上达到最新最好性能,提升规划可靠性
- 根据路况复杂度自动决定是否深入反思,兼顾效率
近期视觉-语言-动作(VLA)模型通过引入推理提升了端到端自动驾驶的可解释性与规划质量。然而,多数方法直接生成最终轨迹,未显式评估其未来后果,限制了在复杂动态环境中的可靠性。为此,我们提出IRRs-Drive(Intend, Reflect, Refine),一种自适应多模态反思框架。该框架首先生成初步文本意图,并通过预测未来语义鸟瞰图(BEV)来预判潜在交互,构建文本+BEV的双模态反思空间,显式建模场景演化,实现对初始意图的严格自纠正。为平衡性能与效率,我们构建面向反思的训练数据,设计自适应反思奖励,使模型根据场景复杂度动态选择推理模式。不同于将推理作为辅助解释,IRRs-Drive将自适应反思机制直接融入规划流程,实现基于场景复杂度的、有依据的决策修正。在NAVSIM基准上,该方法在PDMS和EPDMS指标均达到当前最优表现。大量实验验证了多模态反思框架的有效性及自适应策略的优越性。
原文摘要 · Abstract (English)
Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most existing approaches directly generate the final trajectory without explicitly examining its future consequences, which limits their reliability in complex and dynamic environments. To address this limitation, we propose IRR-Drive (Intend, Reflect, Refine), an adaptive multimodal reflection framework for autonomous driving. Specifically, to tightly couple high-level reasoning with physical constraints, IRR-Drive first generates a preliminary textual intention and anticipates potential interactions by predicting future semantic bird's-eye view (BEV) representations. This dual-modality (Text + BEV) reflection space explicitly models anticipated scene evolution, enabling the model to rigorously self-correct and refine its initial intent before generating the final trajectory. Furthermore, to balance planning performance and computational efficiency, we construct reflection-oriented training data and design an adaptive reflection reward, enabling the model to adaptively select its reasoning mode according to scene complexity. Instead of using reasoning primarily as an auxiliary interpretation, IRR-Drive directly integrates an adaptive reflection mechanism into the planning framework, enabling grounded, decision-aware trajectory correction that is driven by scene complexity. Our method achieves state-of-the-art performance on the NAVSIM benchmark in both PDMS and EPDMS. Extensive experiments demonstrate the effectiveness of our multimodal reflection framework and validate the efficacy of the proposed adaptive reflection strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。