arXiv:2503.07825cs.HCcs.CV2025-03被引 4

用事件相机实现超低功耗手势识别,让智能眼镜更自然好用。

Helios 2.0: A Robust, Ultra-Low Power Gesture Recognition System Optimised for Event-Sensor based Wearables

  • 选小动作手势+仿真数据训练,适应不同用户和环境。
  • 功耗仅6-8毫瓦,准确率超80%,全靠合成数据训练。
  • 适合做可穿戴设备的轻量级视觉交互系统,能省电又精准。

我们提出一种面向可穿戴设备的移动端优化、实时、超低功耗事件相机系统,支持智能眼镜的自然手势控制,显著提升用户体验。尽管计算机视觉中的手势识别已取得进展,但实现直观、跨用户与环境自适应且能耗极低的可穿戴系统仍面临挑战。为此,我们采用精心设计的微手势:拇指在食指上横向滑动(双向)及拇指与食指捏合两次。这些以人为核心的设计利用自然手部动作,无需学习复杂指令即可操作。为应对用户与环境差异,我们开发了新型仿真方法,在无需大量真实数据采集的前提下实现全面领域采样。所提出的功耗优化架构表现优异,在涵盖多样用户与环境的基准数据集上F1得分超过80%。模型在高通骁龙Hexagon DSP上运行时功耗仅为6-8毫瓦;2通道实现准确率超70%,6通道模型所有手势类别均突破80%准确率,全部基于合成数据训练。相比现有最佳方案,准确率提升20%,功耗降低25倍(使用DSP)。该成果推动了超低功耗视觉系统在可穿戴设备中的部署,为无缝人机交互开辟新可能。

原文摘要 · Abstract (English)

We present an advance in wearable technology: a mobile-optimized, real-time, ultra-low-power event camera system that enables natural hand gesture control for smart glasses, dramatically improving user experience. While hand gesture recognition in computer vision has advanced significantly, critical challenges remain in creating systems that are intuitive, adaptable across diverse users and environments, and energy-efficient enough for practical wearable applications. Our approach tackles these challenges through carefully selected microgestures: lateral thumb swipes across the index finger (in both directions) and a double pinch between thumb and index fingertips. These human-centered interactions leverage natural hand movements, ensuring intuitive usability without requiring users to learn complex command sequences. To overcome variability in users and environments, we developed a novel simulation methodology that enables comprehensive domain sampling without extensive real-world data collection. Our power-optimised architecture maintains exceptional performance, achieving F1 scores above 80\% on benchmark datasets featuring diverse users and environments. The resulting models operate at just 6-8 mW when exploiting the Qualcomm Snapdragon Hexagon DSP, with our 2-channel implementation exceeding 70\% F1 accuracy and our 6-channel model surpassing 80\% F1 accuracy across all gesture classes in user studies. These results were achieved using only synthetic training data. This improves on the state-of-the-art for F1 accuracy by 20\% with a power reduction 25x when using DSP. This advancement brings deploying ultra-low-power vision systems in wearable devices closer and opens new possibilities for seamless human-computer interaction.

手势识别事件相机低功耗可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。