arXiv:2604.09658cs.HCcs.CV2026-04被引 1

轻量级眼球手势识别,可在手机端实时运行

TinyGaze: Lightweight Gaze-Gesture Recognition on Commodity Mobile Devices

论文配图:TinyGaze: Lightweight Gaze-Gesture Recognition on Commodity Mobile Devices
图 1 · 摘自论文原文
  • 基于头部/眼部姿态数据构建轻量级时序模型
  • 46千参数模型实现96%手势识别准确率
  • 适合资源受限的移动端眼动交互应用

眼球手势可实现移动设备的免手输入,但实际应用需满足(1)用户易学易记的手势设计,(2)适合在设备端部署的高效识别模型。本文提出一个端到端流程,利用通用ARKit头眼姿态变换,并结合基于学习理论的引导-回忆协议。在小规模可行性研究中(N=4参与者;240次试验;单次会话控制环境),对比紧凑时序模型TinyHAR与更深层基线模型(DeepConvLSTM, SA-HAR)在五分类手势识别和四分类用户识别任务上的表现。TinyHAR在该初步测试中表现优异(手势识别宏F1=0.960;用户识别宏F1=0.997),且仅使用46k参数。模态分析进一步表明,头部姿态动态对移动眼球手势具有高度信息量,凸显了头眼协同作为设计关键点的重要性。尽管样本量小且环境受控限制泛化性,结果仍提示了移动端眼球手势识别的可行方向。

原文摘要 · Abstract (English)

Gaze gestures can provide hands free input on mobile devices, but practical use requires (i) gestures users can learn and recall and (ii) recognition models that are efficient enough for on-device deployment. We present an end-to-end pipeline using commodity ARKit head/eye transforms and a scaffolded guidance-to-recall protocol grounded in learning theory. In a pilot feasibility study (N=4 participants; 240 trials; controlled single-session setting), we benchmark a compact time-series model (TinyHAR) against deeper baselines (DeepConvLSTM, SA-HAR) on 5-way gesture recognition and 4-way user identification. TinyHAR achieves strong performance in this pilot benchmark (Macro F1 = 0.960 for gesture recognition; Macro F1 = 0.997 for user identification) while using only 46k parameters. A modality analysis further indicates that head pose dynamics are highly informative for mobile gaze gestures, highlighting embodied head--eye coordination as a key design consideration. Although the small sample size and controlled setting limit generalizability, these results indicate a potential direction for further investigation into on-device gaze gesture recognition.

眼球追踪轻量模型移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。