通过三元标签一致性降低眼动估计中的不确定因素。
Suppressing Uncertainty in Gaze Estimation
- 引入三元标签一致性度量,融合邻近标签与预测标签评估不确定性。
- 在多个基准上达到当前最优性能,显著提升估计精度。
- 适合处理标注错误或图像质量差的复杂真实场景数据。
眼动估计中的不确定性主要体现在两方面:一是由遮挡、模糊、眼动不一致甚至非人脸图像导致的低质量图像;二是标注过程中标签与实际注视点不匹配引起的错误标签。让这些不确定性参与训练会阻碍模型性能提升。为此,本文提出一种名为SUGE(Suppressing Uncertainty in Gaze Estimation)的有效方法,引入新颖的三元标签一致性度量来估计并抑制不确定性。具体而言,对每个训练样本,我们提出计算一个由邻近样本线性加权投影得到的“邻近标签”,以捕捉图像特征与其对应标签间的相似性关系,并将其与预测伪标签和真实标签结合用于不确定性估计。通过建模这种三元标签一致性,可同时衡量图像与标签的质量,进而通过设计的样本权重调整与标签修正策略,大幅降低劣质图像和错误标签的负面影响。在多个眼动估计基准上的实验结果表明,所提SUGE方法达到当前最优性能。
原文摘要 · Abstract (English)
Uncertainty in gaze estimation manifests in two aspects: 1) low-quality images caused by occlusion, blurriness, inconsistent eye movements, or even non-face images; 2) incorrect labels resulting from the misalignment between the labeled and actual gaze points during the annotation process. Allowing these uncertainties to participate in training hinders the improvement of gaze estimation. To tackle these challenges, in this paper, we propose an effective solution, named Suppressing Uncertainty in Gaze Estimation (SUGE), which introduces a novel triplet-label consistency measurement to estimate and reduce the uncertainties. Specifically, for each training sample, we propose to estimate a novel ``neighboring label'' calculated by a linearly weighted projection from the neighbors to capture the similarity relationship between image features and their corresponding labels, which can be incorporated with the predicted pseudo label and ground-truth label for uncertainty estimation. By modeling such triplet-label consistency, we can measure the qualities of both images and labels, and further largely reduce the negative effects of unqualified images and wrong labels through our designed sample weighting and label correction strategies. Experimental results on the gaze estimation benchmarks indicate that our proposed SUGE achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。