arXiv:2503.12281cs.CVcs.LG2025-03中稿 · 16th ITS European …被引 3

用视觉语言模型提升司机监控系统,零样本识别驾驶状态。

Exploration of VLMs for Driver Monitoring Systems Applications

  • 用VLM直接分析车载视频,无需针对特定任务训练
  • 在司机监控数据集上实现零样本检测,表现优于传统方法
  • 适合快速部署于智能汽车,尤其适合缺乏标注数据的场景

近年来,深度学习模型取得显著进展,尤其是大语言模型(LLMs)和视觉语言模型(VLMs)。这些模型展现出广泛的知识与零样本能力,标志着人工智能进入新阶段,从依赖数据采集和模型训练转向仅通过提示工程即可构建解决方案。尽管该技术已在多个领域应用,包括汽车工业,但其在司机监控系统(DMS)中的研究仍较薄弱。本文提出初步方案,将VLM应用于DMS,使用司机监控数据集评估性能,并讨论其在真实场景中部署的优势与挑战。

原文摘要 · Abstract (English)

In recent years, we have witnessed significant progress in emerging deep learning models, particularly Large Language Models (LLMs) and Vision-Language Models (VLMs). These models have demonstrated promising results, indicating a new era of Artificial Intelligence (AI) that surpasses previous methodologies. Their extensive knowledge and zero-shot capabilities suggest a paradigm shift in developing deep learning solutions, moving from data capturing and algorithm training to just writing appropriate prompts. While the application of these technologies has been explored across various industries, including automotive, there is a notable gap in the scientific literature regarding their use in Driver Monitoring Systems (DMS). This paper presents our initial approach to implementing VLMs in this domain, utilising the Driver Monitoring Dataset to evaluate their performance and discussing their advantages and challenges when implemented in real-world scenarios.

视觉语言模型司机监控零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。