arXiv:2504.09960cs.CV2025-04CVPR被引 4

提升事件相机眼动追踪的鲁棒性与动态适应能力,适合实际部署。

Dual-Path Enhancements in Event-Based Eye Tracking: Augmented Robustness and Adaptive Temporal Modeling

  • 双路径设计:数据增强与混合架构协同提升抗干扰能力。
  • 在3ET+基准上实现1.61的欧氏距离误差,比基线低12%。
  • 适用于AR/VR等实时场景,尤其适合边缘设备运行。

事件相机眼动追踪已成为增强现实与人机交互的关键技术。然而,现有方法在面对突发眼动和环境噪声时表现不佳。基于轻量级时空网络(一种为边缘设备优化的因果架构),本文提出两项关键改进:其一,引入包含时间偏移、空间翻转和事件删除的数据增强流程,显著提升模型鲁棒性,在困难样本上将欧氏距离误差降低12%(1.61 vs. 1.70基线);其二,提出KnightPupil,一种融合EfficientNet-B3空间特征提取、双向GRU上下文建模及线性时变状态空间模块的混合架构,可动态适应稀疏输入与噪声。在3ET+基准上,该框架在CVPR 2025事件相机眼动追踪挑战赛私有测试集上达到1.61的欧氏距离,验证了其在AR/VR系统中实用部署的有效性,并为类脑视觉未来创新提供基础。

原文摘要 · Abstract (English)

Event-based eye tracking has become a pivotal technology for augmented reality and human-computer interaction. Yet, existing methods struggle with real-world challenges such as abrupt eye movements and environmental noise. Building on the efficiency of the Lightweight Spatiotemporal Network-a causal architecture optimized for edge devices-we introduce two key advancements. First, a robust data augmentation pipeline incorporating temporal shift, spatial flip, and event deletion improves model resilience, reducing Euclidean distance error by 12% (1.61 vs. 1.70 baseline) on challenging samples. Second, we propose KnightPupil, a hybrid architecture combining an EfficientNet-B3 backbone for spatial feature extraction, a bidirectional GRU for contextual temporal modeling, and a Linear Time-Varying State-Space Module to adapt to sparse inputs and noise dynamically. Evaluated on the 3ET+ benchmark, our framework achieved 1.61 Euclidean distance on the private test set of the Event-based Eye Tracking Challenge at CVPR 2025, demonstrating its effectiveness for practical deployment in AR/VR systems while providing a foundation for future innovations in neuromorphic vision.

眼动追踪事件相机边缘计算神经形态视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。