arXiv:2507.22041cs.CV2025-07

轻量模型LCN-4在细粒度少样本学习中表现优异,挑战深度网络必优的共识。

Shallow Deep Learning Can Still Excel in Fine-Grained Few-Shot Learning

  • 设计位置感知特征聚类模块,融合空间特征与隐性位置信息。
  • 在三个基准上超越基于ConvNet-4的SOTA,媲美或超过多数ResNet12方法。
  • 适合资源受限场景下追求高效准确的少样本学习应用。

深度学习在细粒度少样本学习(FGFSL)中广泛应用,依赖深层主干网络。然而,浅层网络如ConvNet-4因易提取大量非抽象视觉属性而不受青睐。本文重新评估网络深度与编码少样本实例能力的关系,探索浅层架构是否可达到甚至超越主流深度主干。受原始ConvNet-4启发,提出位置感知星群网络LCN-4,引入先进的位置感知特征聚类模块,能有效编码并整合空间特征融合、特征聚类与隐性特征位置,显著降低整体损失。创新性地提出通用网格位置编码补偿机制,解决普通卷积中位置信息缺失问题;进一步提出通用频域位置嵌入技术,缓解聚类特征中的位置损失。在三个代表性细粒度少样本基准上验证,实验表明LCN-4显著优于基于ConvNet-4的现有最优方法,并达到或超越多数基于ResNet12的方法,证实了其有效性。

原文摘要 · Abstract (English)

Deep learning has witnessed the extensive utilization across a wide spectrum of domains, including fine-grained few-shot learning (FGFSL) which heavily depends on deep backbones. Nonetheless, shallower deep backbones such as ConvNet-4, are not commonly preferred because they're prone to extract a larger quantity of non-abstract visual attributes. In this paper, we initially re-evaluate the relationship between network depth and the ability to fully encode few-shot instances, and delve into whether shallow deep architecture could effectuate comparable or superior performance to mainstream deep backbone. Fueled by the inspiration from vanilla ConvNet-4, we introduce a location-aware constellation network (LCN-4), equipped with a cutting-edge location-aware feature clustering module. This module can proficiently encoder and integrate spatial feature fusion, feature clustering, and recessive feature location, thereby significantly minimizing the overall loss. Specifically, we innovatively put forward a general grid position encoding compensation to effectively address the issue of positional information missing during the feature extraction process of specific ordinary convolutions. Additionally, we further propose a general frequency domain location embedding technique to offset for the location loss in clustering features. We have carried out validation procedures on three representative fine-grained few-shot benchmarks. Relevant experiments have established that LCN-4 notably outperforms the ConvNet-4 based State-of-the-Arts and achieves performance that is on par with or superior to most ResNet12-based methods, confirming the correctness of our conjecture.

少样本学习轻量模型特征聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。