用轻量注意力结构提升边缘设备视频识别效率
CA3D: Convolutional-Attentional 3D Nets for Efficient Video Activity Recognition on the Edge
- 结合卷积与线性复杂度注意力,降低计算开销
- 在多个基准上实现更高准确率,同时保持低计算成本
- 适合对效率和隐私要求高的智能家居与医疗场景
本文提出一种用于视频活动识别的深度学习方法,融合卷积层与线性复杂度注意力机制,显著降低计算开销。我们还引入一种新型量化机制,在训练和推理阶段进一步提升模型效率。该模型在保持强泛化能力的同时,显著降低计算成本,适用于消费级和边缘设备。我们在多个公开可用的视频活动识别基准上验证了模型性能,结果表明其在保持竞争力计算成本的前提下,实现了更高的准确率,为智能家居与智慧医疗等对效率和隐私敏感的应用提供了可行方案。
原文摘要 · Abstract (English)
In this paper, we introduce a deep learning solution for video activity recognition that leverages an innovative combination of convolutional layers with a linear-complexity attention mechanism. Moreover, we introduce a novel quantization mechanism to further improve the efficiency of our model during both training and inference. Our model maintains a reduced computational cost, while preserving robust learning and generalization capabilities. Our approach addresses the issues related to the high computing requirements of current models, with the goal of achieving competitive accuracy on consumer and edge devices, enabling smart home and smart healthcare applications where efficiency and privacy issues are of concern. We experimentally validate our model on different established and publicly available video activity recognition benchmarks, improving accuracy over alternative models at a competitive computing cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。