多视角多模态眼动数据集AA,支持更鲁棒的屏幕注视估计。
AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation
- 8个屏内+2个侧视摄像头同步采集面部图像
- 每样本含多视角人脸与结构化区域裁剪,支持全局与局部特征学习
- 提供独立于被试的评估划分,适合眼动建模与跨主体研究
我们提出AA,一个用于屏幕注视估计的多视角多模态数据集。该数据集通过八个固定在屏幕上的摄像头和两个额外的侧视摄像头,同步采集面部观测,并配以受控注视条件下精确的屏幕空间注视目标。每个样本包含多视角人脸观测及结构化的面部区域裁剪,支持从全局与局部视觉线索进行多模态学习。与现有单视角眼动数据集不同,AA提供了来自屏幕安装和侧向安装视角的多视角覆盖,增强了在视角变化和遮挡情况下的建模鲁棒性。数据集包含独立于被试的评估划分和标准化数据处理流程,支持可复现的眼动估计研究。
原文摘要 · Abstract (English)
We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and two additional side-view cameras, paired with precise screen-space gaze targets collected under controlled fixation conditions. Each sample contains multi-view face observations together with structured facial region crops, enabling multimodal learning from both global and local visual cues. Unlike existing single-view gaze datasets, AA provides multi-view coverage from both screen-mounted and side-mounted perspectives, enabling more robust modeling under viewpoint variation and occlusion. The dataset includes subject-independent evaluation splits and a standardized data processing pipeline to support reproducible research in gaze estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。