用少量人工标注让视觉语言模型描述司机注意力转移原因。
Interpretable Modeling of Driver Attention Shifts with a Vision-Language Model
- 用80个专家标注的注意力变化样本微调视觉语言模型
- 微调后在语义匹配和人类理解度上显著提升
- 适合人因分析与驾驶状态监控场景
司机凝视通常被建模为空间热图,但热图难以解释具体关注对象或注意力转移的意义。本研究探讨在少量人工监督下,能否引导视觉-语言模型生成可解释的注意力转移描述。基于伯克利深度驾驶-注意力数据集中的高变化凝视时刻,对比零样本、单样本及LoRA微调的VLM表现,结果表明:使用80个专家修正的注意力示例进行微调,可显著提升ROUGE-L、METEOR、实体对齐F1与人类对齐分数,优于未引导的VLM输出。研究证明,语言描述能有效补充热图,使驾驶员注意力更易于用于人因分析、驾驶监控审查与情境感知支持。
原文摘要 · Abstract (English)
Driver gaze is commonly modeled as a spatial heatmap, but heatmaps alone are difficult for humans to interpret because they do not explain which road object or region is being monitored or why an attention shift may matter. This study examines whether minimal human-grounded supervision can steer a vision--language model toward interpretable descriptions of driver attention shifts. Using selected high-change gaze moments from the Berkeley DeepDrive-Attention dataset, we compare zero-shot, one-shot, and LoRA fine-tuned VLM conditions against human-refined reference descriptions and expert ratings. Results show that fine-tuning with 80 expert-refined attention examples improves ROUGE-L, METEOR, Entity Alignment F1, and Human Alignment Score relative to unsteered VLM outputs. The findings suggest that language-based descriptions can complement gaze heatmaps by making driver attention more accessible for human-factors analysis, driver-monitoring review, and situation-awareness support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。