arXiv:2511.19629cs.CV2025-11被引 4

用视线数据高效评估第一人称技能水平,省电73倍。

SkillSight: Efficient First-Person Skill Assessment with Gaze

  • 结合视频与视线联合建模,再压缩成仅用视线的轻量模型
  • 在烹饪、音乐、运动三类数据上达到顶尖准确率
  • 适合智能眼镜等低功耗设备,推动真实场景技能学习

智能眼镜的自我中心感知有望改变人们在现实世界中学习新技能的方式,但自动技能评估仍是核心挑战。我们提出 SkillSight,一种从第一人称数据中实现高效技能评估的方法。核心假设是:技能水平不仅体现在动作表现(视频),也体现在执行时的注意力分布(视线)。我们的两阶段框架先联合建模视线与第一人称视频以预测技能水平,再将模型蒸馏为仅依赖视线的轻量学生模型。推理时,仅需输入视线数据,大幅减少功耗,无需持续处理视频。在涵盖烹饪、音乐和体育的三个数据集上的实验首次证明了视线在跨多样化真实场景中对技能理解的重要价值。SkillSight 教师模型达到当前最优性能,而其仅用视线的学生版本在保持高准确率的同时,功耗比现有方法降低73倍。这些结果为野外环境下人工智能支持的技能学习铺平了道路。

原文摘要 · Abstract (English)

Egocentric perception on smart glasses could transform how we learn new skills in the physical world, but automatic skill assessment remains a fundamental technical challenge. We introduce SkillSight for power-efficient skill assessment from first-person data. Central to our approach is the hypothesis that skill level is evident not only in how a person performs an activity (video), but also in how they direct their attention when doing so (gaze). Our two-stage framework first learns to jointly model gaze and egocentric video when predicting skill level, then distills a gaze-only student model. At inference, the student model requires only gaze input, drastically reducing power consumption by eliminating continuous video processing. Experiments on three datasets spanning cooking, music, and sports establish, for the first time, the valuable role of gaze in skill understanding across diverse real-world settings. Our SkillSight teacher model achieves state-of-the-art performance, while our gaze-only student variant maintains high accuracy using 73x less power than competing methods. These results pave the way for in-the-wild AI-supported skill learning.

技能评估视线追踪低功耗智能眼镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。