从单张图片估算司机看广告牌时长,无需眼动设备
Billboard in Focus: Estimating Driver Gaze Duration from a Single Image
- 用YOLO模型检测广告牌,结合DINOv2特征分类
- 单帧识别准确率达68.1%,在BillboardLamac数据集上验证
- 无需人工标注或眼动仪,适合大规模交通行为分析
道路广告牌是户外广告的核心元素,但其存在可能引发驾驶分心和事故风险。本研究提出一种完全自动化的广告牌检测与司机注视时长估算流水线,旨在不依赖人工标注或眼动设备的情况下评估广告牌的视觉吸引力。该流程分为两阶段:(1)基于YOLO的对象检测模型,在Mapillary Vistas上训练,并在BillboardLamac图像上微调,广告牌检测任务达到94% mAP@50;(2)基于检测框位置与DINOv2特征的分类器,实现从单帧图像中估算司机注视时长。实验表明,该方法在考虑单帧图像时,于BillboardLamac数据集上取得68.1%的准确率。结果进一步通过谷歌街景采集的图像得到验证。
原文摘要 · Abstract (English)
Roadside billboards represent a central element of outdoor advertising, yet their presence may contribute to driver distraction and accident risk. This study introduces a fully automated pipeline for billboard detection and driver gaze duration estimation, aiming to evaluate billboard relevance without reliance on manual annotations or eye-tracking devices. Our pipeline operates in two stages: (1) a YOLO-based object detection model trained on Mapillary Vistas and fine-tuned on BillboardLamac images achieved 94% mAP@50 in the billboard detection task (2) a classifier based on the detected bounding box positions and DINOv2 features. The proposed pipeline enables estimation of billboard driver gaze duration from individual frames. We show that our method is able to achieve 68.1% accuracy on BillboardLamac when considering individual frames. These results are further validated using images collected from Google Street View.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。