arXiv:2603.04098cs.CVcs.HC2026-03

用眼神稳定性和瞳孔反应筛选关键帧,提升可穿戴设备的视觉学习效率。

Real Eyes Realize Faster: Gaze Stability and Pupil Novelty for Efficient Egocentric Learning

  • 通过注视稳定性判断画面质量,瞳孔变化捕捉新奇时刻,双标准筛选关键帧。
  • 10%帧率下分类性能媲美全量视频,且信号融合会破坏两者优势。
  • 无需模型推理,实时处理,适合长期运行的智能眼镜等设备使用。

始终开启的可穿戴眼动相机正被广泛用于具身机器人、模仿学习和辅助增强现实,但其视频流充斥冗余与低质帧。在存储与电池受限条件下,选择保留哪些帧与如何学习同等重要。我们发现现代眼动头显提供连续、无需训练的辅助信号,可分解为两个互补维度:注视固定反映视觉稳定性(质量),瞳孔反应反映唤醒相关的新奇时刻(新颖性)。据此提出双准则帧筛选器,先按注视质量过滤帧,再以瞳孔信号对剩余帧排序。在视觉体验数据集(VEDB)上,仅保留10%的帧即达到全量流的分类性能;而简单信号融合反而削弱两者贡献。效果具有任务依赖性:瞳孔排序提升动作识别,而仅靠注视选择已显著优于场景识别,证实两信号发挥着本质不同的作用。该方法无需模型推理,可在采集时实时运行,为持续可用的可穿戴数据筛选提供了可行路径。

原文摘要 · Abstract (English)

Always-on egocentric cameras are increasingly used as demonstrations for embodied robotics, imitation learning, and assistive AR, but the resulting video streams are dominated by redundant and low-quality frames. Under the storage and battery constraints of wearable devices, choosing which frames to keep is as important as how to learn from them. We observe that modern eye-tracking headsets provide a continuous, training-free side channel that decomposes into two complementary axes: gaze fixation captures visual stability (quality), while pupil response captures arousal-linked moments (novelty). We operationalize this insight as a Dual-Criterion Frame Curator that first gates frames by gaze quality and then ranks the survivors by pupil-derived novelty. On the Visual Experience Dataset (VEDB), curated frames at 10% budget match the classification performance of the full stream, and naive signal fusion consistently destroys both contributions. The benefit is task-dependent: pupil ranking improves activity recognition, while gaze-only selection already dominates for scene recognition, confirming that the two signals serve genuinely different roles. Our method requires no model inference and operates at capture time, offering a path toward efficient, always-on egocentric data curation.

眼动追踪数据压缩边缘计算自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。