构建细粒度驾驶行为描述数据集,提升视觉语言模型对驾驶动作的理解能力。
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset

- 基于Drive&Act数据集创建带细粒度自然语言描述的新数据集。
- 微调后模型在DMD数据集上准确率提升至76,优于零样本基线的66。
- 为驾驶监控系统提供可复用的高质量标注数据,适合自动驾驶研究者使用。
理解细微的驾驶行为对构建可靠的驾驶员监控系统至关重要。现有视觉语言模型(VLMs)在通用数据集上训练,难以识别驾驶行为中的细微差别。本文通过创建Drive&Act数据集的详细自然语言版本,解决了这一问题。我们采用基于大语言模型的评分方法评估了三种VLMs在新基准上的表现,发现它们无法可靠生成精确的细粒度驾驶行为描述。基于标注的Drive&Act数据集,我们构建了新的驱动行为描述数据集,用于训练VLMs。在驾驶员监控数据集(DMD)上的跨数据集评估表明,微调后的模型在DMD数据集上具有良好泛化能力。在新数据集上微调的VLM达到76的ACCR分数,优于零样本基线的66。结果表明,使用丰富描述的驾驶行为数据可显著提升模型对驾驶行为的解读能力,同时凸显未来需要更多样化的数据集以支持更广泛的应用。我们的Drive&Act描述数据集和代码将公开发布于GitHub。
原文摘要 · Abstract (English)
Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on general datasets and struggle to recognize fine distinctions in driver behaviors. This paper addresses this limitation by creating a detailed natural language version of the Drive&Act dataset. We evaluate three VLMs on our new benchmark using LLM-based scoring methods. Their performance on the new benchmark shows that they cannot reliably generate accurate fine-grained driver activity descriptions. Based on the labeled Drive&Act dataset we create a new Drive&Act description dataset containing finegrained descriptions to train VLMs on driver activity understanding. Cross dataset evaluation on the Driver Monitoring Dataset (DMD) shows that the VLM fine-tuned on our new Drive&Act description dataset generalizes well to actions in the DMD dataset. The VLM fine-tuned on our Drive&Act description dataset achieves an ACCR score of 76 outperforming the zero-shot VLM baseline with an ACCR score of 66. These findings demonstrate that adapting VLMs with richly described driver actions can significantly improve their ability to interpret driver behavior while also highlighting the need for more diverse datasets to support broader generalization in future applications. Our Drive&Act description dataset and code will be publicly available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。