arXiv:2502.20128cs.CV2025-02中稿 · ACM MM 2025被引 6

用CLIP提升眼神估计精度,新方法让模型更懂眼神的语义差异。

Differential Contrastive Training for Gaze Estimation

  • 通过视觉与语义双分支设计,结合CLIP理解眼神的深层含义。
  • 在四个数据集上实现跨域任务最佳表现,精度显著超越现有方法。
  • 适合做高精度眼神识别、人机交互或通用视觉模型迁移的研究者。

复杂应用场景对精确且泛化能力强的眼神估计方法提出了迫切需求。尽管预训练的CLIP在多种视觉任务中表现优异,但在眼神估计中的潜力尚未充分挖掘。本文提出一种新型微分对比学习策略,借助CLIP增强眼神估计性能。为此,我们构建了包含视觉外观感知分支与语义差异感知分支的DCGaze网络。前者为基本眼神估计网络,集成自适应特征优化单元(AFU)和双头回归器(DGR),有效提取与眼神相关的外观特征;后者基于CLIP文本编码器,揭示眼神间的语义差异,进一步赋予主干网络刻画眼神相关语义信息的能力。在四个具有挑战性的数据集上,针对跨域与同域任务的大量实验验证了本方法的有效性。

原文摘要 · Abstract (English)

The complex application scenarios have raised critical requirements for precise and generalizable gaze estimation methods. Recently, the pre-trained CLIP has achieved remarkable performance on various vision tasks, but its potentials have not been fully exploited in gaze estimation. In this paper, we propose a novel Differential Contrastive Training strategy, which boosts gaze estimation performance with the help of the CLIP. Accordingly, a Differential Contrastive Gaze Estimation network (DCGaze) composed of a Visual Appearance-aware branch and a Semantic Differential-aware branch is introduced. The Visual Appearance-aware branch is essentially a primary gaze estimation network and it incorporates an Adaptive Feature-refinement Unit (AFU) and a Double-head Gaze Regressor (DGR), which both help the primary network to extract informative and gaze-related appearance features. Moreover, the Semantic Difference-aware branch is designed on the basis of the CLIP's text encoder to reveal the semantic difference of gazes. This branch could further empower the Visual Appearance-aware branch with the capability of characterizing the gaze-related semantic information. Extensive experimental results on four challenging datasets over within and cross-domain tasks demonstrate the effectiveness of our DCGaze.The code is available at https://github.com/LinZhang-bjtu/DCGaze.

眼神估计CLIP对比学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。