arXiv:2607.20072cs.CVcs.HC2026-07被引 1

让眼动估计模型学会识别不可靠预测,提升真实场景下的可靠性。

Factor-Informed Uncertainty Distillation for Gaze Estimation

  • 用可解释的图像质量因素指导不确定性建模
  • 在三个数据集上显著提升不确定性的准确性和错误排序相关性
  • 适合需要实时可靠判断的户外眼动系统应用

深度眼动估计在受控环境下表现良好,但在非受限场景中性能下降,需具备拒绝不可靠预测的能力。单次前向传播的不确定性方法(如异方差回归)仅从像素推断不确定性,缺乏输入有效性提示;基于采样的方法虽准确但计算成本高,难以实时使用。本文提出因子感知不确定性蒸馏(FIUD),采用教师-学生框架,将不确定性与可解释的图像质量失败模式对齐。梯度提升教师模型从光照、清晰度、眼部可见性及对称性等因子预测期望眼动误差;神经学生模型通过课程学习和排序监督,将这些信号蒸馏为轻量级单次前向传播的不确定性头。在ETH-XGaze、Gaze360和MPIIFaceGaze(超过30万样本)上,FIUD在不确定性评估、误差排序相关性以及选择性预测方面优于确定性模型和基于采样的基线,尤其在非受限场景中提升最显著。

原文摘要 · Abstract (English)

Deep gaze estimation works well in controlled capture but degrades in unconstrained settings, where systems must reject unreliable predictions. Single-pass uncertainty (e.g., heteroscedastic regression) infers uncertainty from pixels without explicit input-validity cues, while sampling based methods are often too costly for real time use. We propose Factor-Informed Uncertainty Distillation (FIUD), a teacher-student framework that aligns uncertainty with interpretable image-quality failure modes. A gradient-boosting teacher predicts expected gaze error from factors such as illumination, sharpness, eye visibility and symmetry; a neural student distills these signals via curriculum learning and ranking supervision into a lightweight single-pass uncertainty head. Across ETH-XGaze, Gaze360, and MPIIFaceGaze (>300k samples), FIUD improves uncertainty, error rank correlation and selective prediction versus deterministic and sampling-based baselines, with the largest gains in unconstrained settings.

眼动估计不确定性建模轻量化实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。