用语言模型提升视线估计的泛化能力,让模型更懂语义。
LG-Gaze: Learning Geometry-aware Continuous Prompts for Language-Guided Gaze Estimation
- 将视线估计转为视觉-语言对齐任务,利用语义先验增强泛化
- 设计多模态对比回归损失,自适应调整负样本权重
- 引入几何感知插值法,提升视线嵌入精度,适合跨域场景
视线估计模型的泛化能力常受无关因素干扰,尤其在训练数据有限时。现有领域泛化方法因过度依赖值标签回归而效果有限。受预训练视觉-语言模型启发,我们提出一种新框架——语言引导视线估计(LG-Gaze),将视线估计重构为视觉-语言对齐问题。该框架通过提出的多模态对比回归损失,实现视线特征与连续语言特征的对齐,并自适应调整不同负样本的权重。此外,为更好适配视线估计标签,我们设计了几何感知插值方法,生成更精确的视线嵌入。在四个跨域评估任务中,实验验证了该框架的有效性。
原文摘要 · Abstract (English)
The ability of gaze estimation models to generalize is often significantly hindered by various factors unrelated to gaze, especially when the training dataset is limited. Current strategies aim to address this challenge through different domain generalization techniques, yet they have had limited success due to the risk of overfitting when solely relying on value labels for regression. Recent progress in pre-trained vision-language models has motivated us to capitalize on the abundant semantic information available. We propose a novel approach in this paper, reframing the gaze estimation task as a vision-language alignment issue. Our proposed framework, named Language-Guided Gaze Estimation (LG-Gaze), learns continuous and geometry-sensitive features for gaze estimation benefit from the rich prior knowledges of vision-language models. Specifically, LG-Gaze aligns gaze features with continuous linguistic features through our proposed multimodal contrastive regression loss, which customizes adaptive weights for different negative samples. Furthermore, to better adapt to the labels for gaze estimation task, we propose a geometry-aware interpolation method to obtain more precise gaze embeddings. Through extensive experiments, we validate the efficacy of our framework in four different cross-domain evaluation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。