arXiv:2602.01452cs.CVcs.AI2026-02

用 gaze 识别驾驶场景中的物体,大模型表现更优。

Cross-Paradigm Evaluation of Gaze-Based Semantic Object Identification for Intelligent Vehicles

  • 结合眼动与视觉输入,用三类方法识别驾驶视野中的物体
  • 大模型 Qwen2.5-VL-32b 在夜间小物体上准确率超 0.84
  • 传统检测器快但鲁棒性差,大模型理解力强适合安全关键场景

理解驾驶员在驾驶过程中视线关注的位置(即注视行为),对开发下一代智能辅助驾驶系统和提升道路安全至关重要。本文将此问题视为从车辆前视摄像头捕捉的路景中进行语义识别的任务。具体采用三种基于视觉的方法:直接目标检测(YOLOv13)、分割辅助分类(SAM2+EfficientNetV2 与 YOLOv13 对比)以及基于查询的视觉语言模型(VLMs,Qwen2.5-VL-7b 与 Qwen2.5-VL-32b)。结果表明,直接检测(YOLOv13)和 Qwen2.5-VL-32b 显著优于其他方法,宏 F1 分数均超过 0.84。其中,大尺寸 VLM(Qwen2.5-VL-32b)在识别小型、安全关键物体(如交通灯)方面表现出更强鲁棒性,尤其在夜间恶劣条件下表现突出。相反,分割辅助范式因“局部与整体”语义鸿沟导致召回率大幅下降。研究揭示了传统检测器实时性优势与大视觉语言模型更强上下文理解及鲁棒性之间的根本权衡,为未来人机感知型驾驶监控系统的设计提供了关键洞见与实践指导。

原文摘要 · Abstract (English)

Understanding where drivers direct their visual attention during driving, as characterized by gaze behavior, is critical for developing next-generation advanced driver-assistance systems and improving road safety. This paper tackles this challenge as a semantic identification task from the road scenes captured by a vehicle's front-view camera. Specifically, the collocation of gaze points with object semantics is investigated using three distinct vision-based approaches: direct object detection (YOLOv13), segmentation-assisted classification (SAM2 paired with EfficientNetV2 versus YOLOv13), and query-based Vision-Language Models, VLMs (Qwen2.5-VL-7b versus Qwen2.5-VL-32b). The results demonstrate that the direct object detection (YOLOv13) and Qwen2.5-VL-32b significantly outperform other approaches, achieving Macro F1-Scores over 0.84. The large VLM (Qwen2.5-VL-32b), in particular, exhibited superior robustness and performance for identifying small, safety-critical objects such as traffic lights, especially in adverse nighttime conditions. Conversely, the segmentation-assisted paradigm suffers from a "part-versus-whole" semantic gap that led to large failure in recall. The results reveal a fundamental trade-off between the real-time efficiency of traditional detectors and the richer contextual understanding and robustness offered by large VLMs. These findings provide critical insights and practical guidance for the design of future human-aware intelligent driver monitoring systems.

驾驶监控视觉语言模型眼动识别智能汽车

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。