arXiv:2509.18692cs.CV2025-09被引 2

轻量级视觉变换器提升食物图像分类效率与精度

Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification

  • 采用窗口与空间注意力机制降低计算开销
  • 在Food-101和Vireo Food-172上分别达95.24%和94.33%准确率
  • 适合边缘设备部署,兼顾速度与识别效果

随着社会发展与科技进步,食品产业对生产质量与效率的要求日益提高。食物图像分类在自动化质检、食品安全监管及智能农业中具有关键作用。然而,视觉变换器模型参数量大、计算复杂度高,限制了其应用。为此,我们提出一种轻量级食物图像分类算法,融合窗口多头注意力机制(WMHAM)与空间注意力机制(SAM)。WMHAM通过高效窗口划分捕获局部与全局上下文特征,降低计算成本;SAM则自适应强调关键空间区域,增强特征判别力。在Food-101与Vireo Food-172数据集上的实验表明,模型准确率分别达到95.24%和94.33%,显著减少参数量与浮点运算次数。结果证明该方法在计算效率与分类性能间实现良好平衡,适用于资源受限环境部署。

原文摘要 · Abstract (English)

With the rapid development of society and continuous advances in science and technology, the food industry increasingly demands higher production quality and efficiency. Food image classification plays a vital role in enabling automated quality control on production lines, supporting food safety supervision, and promoting intelligent agricultural production. However, this task faces challenges due to the large number of parameters and high computational complexity of Vision Transformer models. To address these issues, we propose a lightweight food image classification algorithm that integrates a Window Multi-Head Attention Mechanism (WMHAM) and a Spatial Attention Mechanism (SAM). The WMHAM reduces computational cost by capturing local and global contextual features through efficient window partitioning, while the SAM adaptively emphasizes key spatial regions to improve discriminative feature representation. Experiments conducted on the Food-101 and Vireo Food-172 datasets demonstrate that our model achieves accuracies of 95.24% and 94.33%, respectively, while significantly reducing parameters and FLOPs compared with baseline methods. These results confirm that the proposed approach achieves an effective balance between computational efficiency and classification performance, making it well-suited for deployment in resource-constrained environments.

视觉变换器图像分类轻量化食物识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。