arXiv:2603.19588cs.HCcs.CV2026-03被引 1

利用屏幕内容知识提升设备端眼动追踪精度

HiFiGaze: Improving Eye Tracking Accuracy Using Screen Content Knowledge

论文配图:HiFiGaze: Improving Eye Tracking Accuracy Using Screen Content Knowledge
图 1 · 摘自论文原文
  • 结合屏幕显示内容识别眼内屏幕反光区域
  • 相比基线模型平均误差降低8%
  • 底部摄像头配置可再提升10%-20%精度

我们提出一种新型高精度眼动追踪方法,适用于智能手机、笔记本等消费级计算设备。随着高端设备摄像头分辨率已达4K及以上,现可捕捉用户眼中屏幕的二维反射图像。然而,由于屏幕内容千变万化,仅靠反射图像难以实现精准追踪。关键在于,设备本身知道屏幕上显示的内容,本研究利用该信息实现对反射区域的鲁棒分割,其位置与大小编码了用户相对于屏幕的注视目标。我们探索多种策略并开展用户实验评估性能,最佳模型相比基线外观模型将平均追踪误差降低约8%。补充研究表明,若眼动追踪摄像头位于设备底部,还可额外提升10%-20%精度。

原文摘要 · Abstract (English)

We present a new and accurate approach for gaze estimation on consumer computing devices. We take advantage of continued strides in the quality of user-facing cameras found in e.g., smartphones, laptops, and desktops - 4K or greater in high-end devices - such that it is now possible to capture the 2D reflection of a device's screen in the user's eyes. This alone is insufficient for accurate gaze tracking due to the near-infinite variety of screen content. Crucially, however, the device knows what is being displayed on its own screen - in this work, we show this information allows for robust segmentation of the reflection, the location and size of which encodes the user's screen-relative gaze target. We explore several strategies to leverage this useful signal, quantifying performance in a user study. Our best performing model reduces mean tracking error by ~8% compared to a baseline appearance-based model. A supplemental study reveals an additional 10-20% improvement if the gaze-tracking camera is located at the bottom of the device.

眼动追踪视觉感知人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。