通过时空建模提升视频眼动估计精度,无需特定人群适应
Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
- 用卷积+通道注意力与自注意力融合眼脸特征,构建时空序列
- 在EVE数据集上达到当前最佳性能,无需个体适配也有效
- 保留帧内空间上下文比过早池化更关键,适合通用摄像头应用
基于视频的眼动估计需捕捉人类眼神的固有时间动态。由于模型需同时建模帧内空间关系与帧间时序关系,性能受限于单帧及多帧间的特征表示。本文提出时空眼动网络(ST-Gaze),结合CNN主干、专用通道注意力与自注意力模块,优化融合眼区与面部特征。融合特征被视作空间序列,以捕捉帧内上下文,并沿时间传播以建模帧间动态。在EVE数据集上的实验表明,ST-Gaze在有无个体适配条件下均达到最先进水平。消融实验进一步揭示:通过时空循环保留并建模帧内空间上下文,显著优于早期空间池化。结果为使用常见摄像头实现更鲁棒的视频眼动估计提供了新路径。
原文摘要 · Abstract (English)
Video-based gaze estimation methods aim to capture the inherently temporal dynamics of human eye gaze from multiple image frames. However, since models must capture both spatial and temporal relationships, performance is limited by the feature representations within a frame but also between multiple frames. We propose the Spatio-Temporal Gaze Network (ST-Gaze), a model that combines a CNN backbone with dedicated channel attention and self-attention modules to fuse eye and face features optimally. The fused features are then treated as a spatial sequence, allowing for the capture of an intra-frame context, which is then propagated through time to model inter-frame dynamics. We evaluated our method on the EVE dataset and show that ST-Gaze achieves state-of-the-art performance both with and without person-specific adaptation. Additionally, our ablation study provides further insights into the model performance, showing that preserving and modelling intra-frame spatial context with our spatio-temporal recurrence is fundamentally superior to premature spatial pooling. As such, our results pave the way towards more robust video-based gaze estimation using commonly available cameras.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。