为多模态大模型在自动驾驶中的场景理解能力设计评估框架
A Framework for a Capability-driven Evaluation of Scenario Understanding for Multimodal Large Language Models in Autonomous Driving
- 按语义、空间、时间、物理四维度构建评估体系
- 涵盖感知、推理、决策等多层任务与模态
- 适合评估自动驾驶中语言驱动的智能系统
多模态大语言模型(MLLMs)有望通过融合通用世界知识与情境化语言引导,提升自动驾驶能力。尽管其在孤立概念验证应用中表现良好,但当前评估仅聚焦于感知、推理或规划的单一维度。为充分挖掘其潜力,本文提出一个面向自动驾驶的、以能力为导向的MLLM评估框架。该框架从自动驾驶系统需求、人类驾驶认知和语言推理出发,将场景理解划分为语义、空间、时间、物理四个核心能力维度,并进一步组织为上下文层级、处理模态与下游任务(如语言交互与决策)。通过两个典型交通场景的分析,验证了该框架在真实驾驶情境中的适用性,为系统评估MLLM在自动驾驶中的场景理解能力提供了基础。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) hold the potential to enhance autonomous driving by combining domain-independent world knowledge with context-specific language guidance. Their integration into autonomous driving systems shows promising results in isolated proof-of-concept applications, while their performance is evaluated on selective singular aspects of perception, reasoning, or planning. To leverage their full potential a systematic framework for evaluating MLLMs in the context of autonomous driving is required. This paper proposes a holistic framework for a capability-driven evaluation of MLLMs in autonomous driving. The framework structures scenario understanding along the four core capability dimensions semantic, spatial, temporal, and physical. They are derived from the general requirements of autonomous driving systems, human driver cognition, and language-based reasoning. It further organises the domain into context layers, processing modalities, and downstream tasks such as language-based interaction and decision-making. To illustrate the framework's applicability, two exemplary traffic scenarios are analysed, grounding the proposed dimensions in realistic driving situations. The framework provides a foundation for the structured evaluation of MLLMs' potential for scenario understanding in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。