arXiv:2409.11051cs.CV2024-09ECCV被引 3

通过下采样适配器实现高效超细粒度图像识别

Down-Sampling Inter-Layer Adapter for Parameter and Computation Efficient Ultra-Fine-Grained Image Recognition

  • 在冻结主干网络前提下,仅微调少量下采样适配模块
  • 相比当前方法参数减少123倍,浮点运算量降低30%
  • 平均准确率提升至少6.8%,适合资源受限场景

超细粒度图像识别(UFGIR)旨在区分同一物种内差异极小的类别(如不同品种),其挑战在于每类样本稀缺。本文提出一种基于下采样跨层适配器的新方法,在参数高效设置下冻结主干网络,仅微调少量附加模块。通过引入双分支下采样机制,显著降低参数量与浮点运算量(FLOPs)。在十个数据集上的实验证明,该方法在参数效率与精度之间取得优异平衡,相比现有参数高效方法平均准确率提升至少6.8%,可训练参数减少至少123倍,平均浮点运算量降低30%,具备在资源受限环境中的实际应用潜力。

原文摘要 · Abstract (English)

Ultra-fine-grained image recognition (UFGIR) categorizes objects with extremely small differences between classes, such as distinguishing between cultivars within the same species, as opposed to species-level classification in fine-grained image recognition (FGIR). The difficulty of this task is exacerbated due to the scarcity of samples per category. To tackle these challenges we introduce a novel approach employing down-sampling inter-layer adapters in a parameter-efficient setting, where the backbone parameters are frozen and we only fine-tune a small set of additional modules. By integrating dual-branch down-sampling, we significantly reduce the number of parameters and floating-point operations (FLOPs) required, making our method highly efficient. Comprehensive experiments on ten datasets demonstrate that our approach obtains outstanding accuracy-cost performance, highlighting its potential for practical applications in resource-constrained environments. In particular, our method increases the average accuracy by at least 6.8\% compared to other methods in the parameter-efficient setting while requiring at least 123x less trainable parameters compared to current state-of-the-art UFGIR methods and reducing the FLOPs by 30\% in average compared to other methods.

超细粒度参数高效下采样图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。