arXiv:2412.08074cs.CVcs.LG2024-12被引 1

轻量级模型EM-Net用期望最大化算法提升注视估计精度

EM-Net: Gaze Estimation with Expectation Maximization Algorithm

  • 结合深度学习与期望最大化算法,设计全局注意力机制提取关键特征
  • 仅用50%数据训练,对三个数据集精度提升超2%
  • 适合资源受限场景,对噪声干扰有强鲁棒性

近年来,注视估计技术的准确率持续提升,但现有方法通常依赖大规模数据集或大模型,导致计算资源需求高。针对此问题,本文提出一种基于深度学习与期望最大化(Expectation Maximization)算法的轻量级注视估计模型EM-Net。首先,在模型中引入全局注意力机制(GAM),以增强对注视相关特征的全局依赖捕捉能力,从而提升性能;其次,通过EM模块学习分层特征表示,增强了模型的泛化能力,降低对样本数量的需求。实验表明,在仅使用50%训练数据的前提下,EM-Net在Gaze360、MPIIFaceGaze和RT-Gene数据集上分别较GazeNAS-ETH提升2.2%、2.02%和2.03%。同时,该模型在高斯噪声干扰下仍表现出良好鲁棒性。

原文摘要 · Abstract (English)

In recent years, the accuracy of gaze estimation techniques has gradually improved, but existing methods often rely on large datasets or large models to improve performance, which leads to high demands on computational resources. In terms of this issue, this paper proposes a lightweight gaze estimation model EM-Net based on deep learning and traditional machine learning algorithms Expectation Maximization algorithm. First, the proposed Global Attention Mechanism(GAM) is added to extract features related to gaze estimation to improve the model's ability to capture global dependencies and thus improve its performance. Second, by learning hierarchical feature representations through the EM module, the model has strong generalization ability, which reduces the need for sample size. Experiments have confirmed that, on the premise of using only 50% of the training data, EM-Net improves the performance of Gaze360, MPIIFaceGaze, and RT-Gene datasets by 2.2%, 2.02%, and 2.03%, respectively, compared with GazeNAS-ETH. It also shows good robustness in the face of Gaussian noise interference.

注视估计轻量模型EM算法低数据需求

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。