用场景注意力提升驾驶员视线估计精度,实测误差降低23.5%。
SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze

- 融合面部与场景图像,通过注意力机制建模驾驶员视线意图。
- 在UD-FSG数据集上达到104.73像素均值误差,较当前最优降低23.5%。
- 尤其在场景边缘区域表现更优,适合真实驾驶环境应用。
驾驶员视线估计对理解其对周围交通的感知至关重要。现有模型仅依赖面部信息预测视线点(PoG)或3D注视方向。本文提出基准数据集Urban Driving-Face Scene Gaze(UD-FSG),包含同步的驾驶员面部与交通场景图像。场景图像提供周围交通线索,有助于提升注视估计性能。我们提出SGAP-Gaze模型,显式融合场景图像进行视线估计,整合驾驶员面部、眼睛、虹膜及场景上下文信息。首先,从面部模态提取特征并融合生成注视意图向量;其次,基于Transformer的注意力机制融合面部与场景特征,在空间场景网格上计算注意力得分以获得PoG。在UD-FSG数据集上,该模型实现104.73像素均值误差,在LBW数据集上为63.48像素,相较当前最优模型降低23.5%。空间像素分布分析显示,无论在场景内还是外区域,本方法均持续优于现有方法,尤其在稀有但关键的边缘区域表现突出。结果验证了多模态线索结合场景感知注意力在真实驾驶环境中的有效性。
原文摘要 · Abstract (English)
Driver gaze estimation is essential for understanding the driver's situational awareness of surrounding traffic. Existing gaze estimation models use driver facial information to predict the Point-of-Gaze (PoG) or the 3D gaze direction vector. We propose a benchmark dataset, Urban Driving-Face Scene Gaze (UD-FSG), comprising synchronized driver-face and traffic-scene images. The scene images provide cues about surrounding traffic, which can help improve the gaze estimation model, along with the face images. We propose SGAP-Gaze, Scene-Grid Attention based Point-of-Gaze estimation network, trained and tested on our UD-FSG dataset, which explicitly incorporates the scene images into the gaze estimation modelling. The gaze estimation network integrates driver face, eye, iris, and scene contextual information. First, the extracted features from facial modalities are fused to form a gaze intent vector. Then, attention scores are computed over the spatial scene grid using a Transformer-based attention mechanism fusing face and scene image features to obtain the PoG. The proposed SGAP-Gaze model achieves a mean pixel error of 104.73 on the UD-FSG dataset and 63.48 on LBW dataset, achieving a 23.5% reduction in mean pixel error compared to state-of-the-art driver gaze estimation models. The spatial pixel distribution analysis shows that SGAP-Gaze consistently achieves lower mean pixel error than existing methods across all spatial ranges, including the outer regions of the scene, which are rare but critical for understanding driver attention. These results highlight the effectiveness of integrating multi-modal gaze cues with scene-aware attention for a robust driver PoG estimation model in real-world driving environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。