arXiv:2507.18447cs.CV2025-07被引 5

构建驾驶行为理解新基准,提升大模型对驾驶意图的解释能力。

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior

  • 设计双模块数据集PDB-X与PDB-QA,从外部视角推断内部驾驶行为。
  • 在脑力车意向预测任务中提升12.5%,问答零样本性能提高73.2%。
  • 适用于自动驾驶系统个性化优化,适合研究多模态推理与智能驾驶者。

理解驾驶员行为与意图对风险评估和事故预防至关重要。安全与辅助系统可针对个体驾驶行为定制,显著提升有效性。然而,现有数据集在基于外部视觉证据描述和解释一般车辆运动方面存在局限。本文提出PDB-Eval基准,用于精细化理解个性化驾驶行为,并对大得多模态模型(MLLMs)进行驾驶理解与推理对齐。该基准包含两个核心组件:PDB-X用于评估模型对时间序列驾驶场景的理解;PDB-QA作为视觉解释问答任务,用于指导MLLM指令微调。作为生成模型通用学习任务,PDB-QA可在不损害泛化能力的前提下弥合领域差距。实验表明,对细粒度描述与解释进行微调,可有效缩小模型与驾驶领域的差距,使问答任务零样本性能最高提升73.2%。进一步评估在Brain4Cars意图预测与AIDE识别任务中的表现,微调后模型在转向意图预测上提升12.5%,在所有AIDE任务中均实现最高11.0%的性能增益。

原文摘要 · Abstract (English)

Understanding a driver's behavior and intentions is important for potential risk assessment and early accident prevention. Safety and driver assistance systems can be tailored to individual drivers' behavior, significantly enhancing their effectiveness. However, existing datasets are limited in describing and explaining general vehicle movements based on external visual evidence. This paper introduces a benchmark, PDB-Eval, for a detailed understanding of Personalized Driver Behavior, and aligning Large Multimodal Models (MLLMs) with driving comprehension and reasoning. Our benchmark consists of two main components, PDB-X and PDB-QA. PDB-X can evaluate MLLMs' understanding of temporal driving scenes. Our dataset is designed to find valid visual evidence from the external view to explain the driver's behavior from the internal view. To align MLLMs' reasoning abilities with driving tasks, we propose PDB-QA as a visual explanation question-answering task for MLLM instruction fine-tuning. As a generic learning task for generative models like MLLMs, PDB-QA can bridge the domain gap without harming MLLMs' generalizability. Our evaluation indicates that fine-tuning MLLMs on fine-grained descriptions and explanations can effectively bridge the gap between MLLMs and the driving domain, which improves zero-shot performance on question-answering tasks by up to 73.2%. We further evaluate the MLLMs fine-tuned on PDB-X in Brain4Cars' intention prediction and AIDE's recognition tasks. We observe up to 12.5% performance improvements on the turn intention prediction task in Brain4Cars, and consistent performance improvements up to 11.0% on all tasks in AIDE.

驾驶行为多模态模型意图预测评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。