arXiv:2506.01757cs.CVcs.AI2025-06中稿 · as an extended abs…

用多模态数据降低可穿戴设备算力需求,实现高效实时动作识别。

Efficient Egocentric Action Recognition with Multimodal Data

  • 融合低频视觉与高频手势数据,优化计算效率。
  • 采样率降低后仍保持高识别准确率,CPU占用减少3倍。
  • 适合资源受限的可穿戴设备实时动作识别场景。

可穿戴XR设备的普及为第一人称动作识别(EAR)系统带来了新机遇,有助于深化对人类行为的理解和情境感知。然而,由于便携性、电池续航与计算资源之间的固有权衡,实现实时算法部署面临挑战。本文系统分析了不同输入模态(RGB视频与3D手部姿态)在不同采样频率下对第一人称动作识别性能及CPU使用率的影响。通过探索多种配置,全面刻画了准确率与计算效率间的权衡关系。研究发现,当以更高频率输入3D手部姿态数据来补充较低采样率的RGB帧时,可在显著降低CPU消耗的同时保持高识别精度。值得注意的是,识别性能几乎无损的情况下,CPU使用量最高可降低3倍。这表明多模态输入策略是实现高效、实时第一人称动作识别在XR设备上的可行路径。

原文摘要 · Abstract (English)

The increasing availability of wearable XR devices opens new perspectives for Egocentric Action Recognition (EAR) systems, which can provide deeper human understanding and situation awareness. However, deploying real-time algorithms on these devices can be challenging due to the inherent trade-offs between portability, battery life, and computational resources. In this work, we systematically analyze the impact of sampling frequency across different input modalities - RGB video and 3D hand pose - on egocentric action recognition performance and CPU usage. By exploring a range of configurations, we provide a comprehensive characterization of the trade-offs between accuracy and computational efficiency. Our findings reveal that reducing the sampling rate of RGB frames, when complemented with higher-frequency 3D hand pose input, can preserve high accuracy while significantly lowering CPU demands. Notably, we observe up to a 3x reduction in CPU usage with minimal to no loss in recognition performance. This highlights the potential of multimodal input strategies as a viable approach to achieving efficient, real-time EAR on XR devices.

动作识别多模态边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。