arXiv:2509.11165cs.CV2025-09

用好奇心驱动学习,让自动驾驶模型在复杂路况中自主提炼经验规律。

Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning

  • 不依赖检索,直接训练出可泛化的交通案例空间
  • 在真实数据集上达83.1%准确率,显著提升长尾场景推理能力
  • 适合研究自动驾驶决策与多模态知识融合的开发者

为实现安全可靠的自动驾驶,决策系统需有效利用过往经验应对交通场景的长尾问题。案例推理(CBR)为此提供自然范式,通过复用历史案例解决方案。然而,在复杂动态交通环境中,传统CBR方法难以在不确定性下有效抽象与适应知识。尽管多模态大语言模型(MLLM)具备强大感知与语言能力,其推理常依赖经验模式匹配,导致分布外与长尾场景鲁棒性不足。本文提出Traffic-MLLM,一种无需检索的神经化案例建模框架,用于多模态交通推理。训练阶段直接学习结构化、可泛化的案例空间,而非推理时检索案例。我们整合动态交通视频与大规模静态视觉问答数据,构建多源案例库作为统一训练基础,以学习结构化案例表征。为进一步提升知识边界处的表示质量,引入基于随机网络蒸馏(RND)的好奇心驱动优化机制,促使模型内化跨案例结构规律,而非表面相关性。在SUTD-TrafficQA与DriveQA基准测试中,模型在动态推理、规则理解与跨域迁移方面均取得一致提升:在SUTD-TrafficQA上达50.8%准确率,在基于CARLA的DriveQA子集上达74.8%,在真实世界Mapillary子集上达83.1%。结果表明,表示层的案例空间优化可作为可扩展多模态案例适配的有效替代方案。

原文摘要 · Abstract (English)

For safe and robust autonomous driving, decision-making systems must effectively leverage past experiences to handle the inherent long-tail of traffic scenarios. Case-Based Reasoning (CBR) provides a natural paradigm for this by adapting solutions from prior cases. However, in complex and dynamic traffic environments, traditional CBR methods struggle to effectively abstract and adapt knowledge under uncertainty. Meanwhile, although multimodal large language models (MLLMs) exhibit strong perceptual and linguistic capabilities, their reasoning behavior often relies on empirical pattern fitting, limiting robustness under distribution shift and long-tail scenarios. We propose Traffic-MLLM, a retrieval-free neural case modeling framework for multimodal traffic reasoning. Instead of performing explicit case retrieval at inference time, Traffic-MLLM learns a structured and generalizable case space directly during training. To support this learning process, we construct a multi-source case base by integrating dynamic traffic videos and large-scale static visual question-answering data, serving as a unified training substrate for learning structured case representations. To further improve representation quality near knowledge boundaries, we introduce a curiosity-driven refinement mechanism based on Random Network Distillation (RND), encouraging the model to internalize cross-case structural regularities rather than surface correlations. Experiments on the SUTD-TrafficQA and DriveQA benchmarks demonstrate consistent improvements in dynamic reasoning, regulatory understanding, and cross-domain transfer. Traffic-MLLM achieves 50.8% accuracy on SUTD-TrafficQA, 74.8% on the CARLA-based DriveQA split, and 83.1% on the real-world Mapillary split, indicating that representation-level case-space refinement provides an effective alternative to explicit retrieval for scalable multimodal case adaptation.

自动驾驶案例推理多模态好奇心机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。