arXiv:2504.11160cs.CVcs.AI2025-04被引 2

通过解耦面部特征与多尺度注意力提升注视估计精度

DMAGaze: Gaze Estimation Based on Feature Disentanglement and Multi-Scale Attention

  • 用双分支掩码解耦器分离面部中的注视相关与无关信息
  • 在两个公开数据集上达到当前最优性能,显著降低干扰影响
  • 适合需要高精度注视追踪的交互系统开发人员

注视估计旨在预测视线方向,常受人脸图像中复杂非注视相关信息干扰。本文提出DMAGaze框架,从三个维度利用面部图像信息:解耦后的全局注视相关特征(来自面部图像)、局部眼部特征(从裁剪的眼部区域提取)以及头部姿态估计特征,以提升整体性能。首先,设计一种基于连续掩码的解耦器,通过分别重建眼部与非眼部区域,实现双分支解耦目标,精准分离注视相关与无关信息。其次,引入新型级联注意力模块MS-GLAM,通过定制化级联结构,在多尺度下有效聚焦全局与局部信息,进一步增强解耦器输出的信息表达。最后,将上半脸分支解耦出的全局注视相关特征,结合头部姿态和局部眼部特征,送入检测头完成高精度注视估计。所提方法在两个主流公开数据集上经过充分验证,达到当前最优性能。

原文摘要 · Abstract (English)

Gaze estimation, which predicts gaze direction, commonly faces the challenge of interference from complex gaze-irrelevant information in face images. In this work, we propose DMAGaze, a novel gaze estimation framework that exploits information from facial images in three aspects: gaze-relevant global features (disentangled from facial image), local eye features (extracted from cropped eye patch), and head pose estimation features, to improve overall performance. Firstly, we design a new continuous mask-based Disentangler to accurately disentangle gaze-relevant and gaze-irrelevant information in facial images by achieving the dual-branch disentanglement goal through separately reconstructing the eye and non-eye regions. Furthermore, we introduce a new cascaded attention module named Multi-Scale Global Local Attention Module (MS-GLAM). Through a customized cascaded attention structure, it effectively focuses on global and local information at multiple scales, further enhancing the information from the Disentangler. Finally, the global gaze-relevant features disentangled by the upper face branch, combined with head pose and local eye features, are passed through the detection head for high-precision gaze estimation. Our proposed DMAGaze has been extensively validated on two mainstream public datasets, achieving state-of-the-art performance.

注视估计特征解耦多尺度注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。