arXiv:2509.20843cs.RO2025-09被引 11

让自动驾驶模型学会从记忆中调用经验,应对复杂突发路况。

MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases

  • 用记忆检索+动态工具箱实现闭环推理,提升环境理解能力
  • 在NAVIM上达88.3的PDMS,规划准确率达82.6%
  • 适合关注鲁棒性与零样本泛化的自动驾驶研究者

视觉语言模型在端到端自动驾驶中展现出巨大潜力,但其在分布外(OOD)场景下的脆弱性仍限制实际部署。本文提出MTRDrive框架,通过融合过程化驾驶经验与动态工具集,构建记忆-工具协同推理机制,增强模型的泛化与主动决策能力。该框架采用闭环系统设计,结合记忆驱动的经验检索与动态工具使用,显著提升环境交互与推理效率。同时,我们构建了基于复杂道路施工场景的新基准Roadwork-VLM,用于评估零样本泛化能力。大量实验表明,3B参数的MTRDrive在公开NAVSIM基准上无链式思维时取得88.3的PDMS,高阶规划指标达79.8%,规划准确率82.6%;在新基准上零样本测试亦达80.2的驾驶指标,证明其在未见场景中的强鲁棒推理能力,推动自动驾驶向更安全可靠方向发展。

原文摘要 · Abstract (English)

Vision-Language Models(VLMs) have demonstrated significant potential for end-to-end autonomous driving, yet a substantial gap remains between their current capabilities and the reliability necessary for real-world deployment. A critical challenge is their fragility, characterized by hallucinations and poor generalization in out-of-distribution (OOD) scenarios. To bridge this gap, we introduce MTRDrive, a novel framework that integrates procedural driving experiences with a dynamic toolkit to enhance generalization and proactive decision-making. MTRDrive addresses these limitations through a closed-loop system that combines a memory-based experience retrieval mechanism with dynamic toolkits. This synergy enables the model to interact more effectively with its environment, improving both reasoning and decision-making capabilities with the help of our memory-tool synergistic reasoning. Additionally, we introduce a new benchmark based on complex Roadwork construction scenarios to rigorously evaluate zero-shot generalization. Extensive experiments demonstrate the superior effectiveness of our approach. On the public NAVSIM benchmark, our 3B-parameter MTRDrive model achieves an exceptional PDMS of 88.3 without chain-of-thought and sets a state-of-the-art performance bar on high-level planning, with a driving metric score of 79.8\% and a planning accuracy of 82.6\%. Rigorous zero-shot evaluation on the new Roadwork-VLM benchmark shows a strong ability to reason robustly in unseen scenarios, achieving a driving metric score of 80.2\%. These results highlight MTRDrive's potential to advance autonomous driving toward safer and more reliable systems.

自动驾驶视觉语言模型记忆推理零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。