arXiv:2411.07834cs.CV2024-11被引 3

用稀疏激活的视觉专家模型,在边缘设备上高效识别鸟类。

Towards Vision Mixture of Experts for Wildlife Monitoring on the Edge

  • 按图像块动态选择专家网络,减少计算量
  • 参数量减4倍,准确率仅降1%
  • 适合资源受限的野外动物监测场景

物联网传感器在工业、消费和遥感领域的爆发式增长,带来了对海量数据传输与分析的计算基础设施需求。与此同时,可持续计算成为关注焦点。为此,学术界正致力于降低深度学习算法的计算开销,尤其在边缘端实现低延迟推理与隐私保护。当前的TinyML研究强调减少通信带宽与云端存储成本。理想的解决方案应能处理时间序列、音频、卫星图像和视频等多模态数据流,以提升模型细粒度判别能力。近期基于数据驱动的条件计算方法已在图文多模态模型中取得进展,实现了跨模态参数共享。受此启发,本文首次将每块条件计算应用于移动视觉变压器(仅视觉),为未来单塔多模态边缘模型铺路。我们在Cornell Sap Sucker Woods 60(SSW60)数据集上评估,相比MobileViTV2-1.0,参数量减少4倍,iNaturalist '21鸟类测试集上准确率仅下降1%。

原文摘要 · Abstract (English)

The explosion of IoT sensors in industrial, consumer and remote sensing use cases has come with unprecedented demand for computing infrastructure to transmit and to analyze petabytes of data. Concurrently, the world is slowly shifting its focus towards more sustainable computing. For these reasons, there has been a recent effort to reduce the footprint of related computing infrastructure, especially by deep learning algorithms, for advanced insight generation. The `TinyML' community is actively proposing methods to save communication bandwidth and excessive cloud storage costs while reducing algorithm inference latency and promoting data privacy. Such proposed approaches should ideally process multiple types of data, including time series, audio, satellite images, and video, near the network edge as multiple data streams has been shown to improve the discriminative ability of learning algorithms, especially for generating fine grained results. Incidentally, there has been recent work on data driven conditional computation of subnetworks that has shown real progress in using a single model to share parameters among very different types of inputs such as images and text, reducing the computation requirement of multi-tower multimodal networks. Inspired by such line of work, we explore similar per patch conditional computation for the first time for mobile vision transformers (vision only case), that will eventually be used for single-tower multimodal edge models. We evaluate the model on Cornell Sap Sucker Woods 60, a fine grained bird species discrimination dataset. Our initial experiments uses $4X$ fewer parameters compared to MobileViTV2-1.0 with a $1$% accuracy drop on the iNaturalist '21 birds test data provided as part of the SSW60 dataset.

边缘计算视觉专家TinyML鸟类识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。