arXiv:2604.14958cs.CV2026-04

通过频域与空间域双子空间融合,提升少样本细粒度分类的稳定性与精度

Frequency-Enhanced Dual-Subspace Networks for Few-Shot Fine-Grained Image Classification

论文配图:Frequency-Enhanced Dual-Subspace Networks for Few-Shot Fine-Grained Image Classification
图 1 · 摘自论文原文
  • 利用DCT和低通滤波分离图像高频噪声与低频结构特征
  • 构建独立纹理与结构子空间,动态加权融合两类特征距离
  • 在多个数据集上超越现有方法,适合高精度细粒度识别场景

少样本细粒度图像分类旨在用极少标注样本识别视觉差异微小的子类别。现有基于度量学习的方法通常仅依赖空间域特征,易受纹理偏差影响,将关键结构细节与高频背景噪声混淆。同时,缺乏跨视角几何约束,单视图度量在少样本条件下易过拟合噪声,导致结构不稳定。为此,本文提出频域增强双子空间网络(FEDSNet)。该方法使用离散余弦变换(DCT)与低通滤波机制,显式分离空间特征中的低频全局结构成分,抑制背景干扰;通过截断奇异值分解(SVD)为纹理与结构特征分别构建独立的低秩线性子空间,并设计自适应门控机制,动态融合两类子空间的投影距离。该策略利用频域子空间的结构稳定性,防止空间子空间过拟合背景特征。在四个基准数据集(CUB-200-2011、Stanford Cars、Stanford Dogs、FGVC-Aircraft)上的大量实验表明,FEDSNet表现出优异的分类性能与鲁棒性,相较现有度量学习算法取得具有竞争力的结果。复杂度分析进一步证实,所提网络在高精度与计算效率间实现良好平衡,为少样本细粒度视觉识别提供了一种有效新范式。

原文摘要 · Abstract (English)

Few-shot fine-grained image classification aims to recognize subcategories with high visual similarity using only a limited number of annotated samples. Existing metric learning-based methods typically rely solely on spatial domain features. Confined to this single perspective, models inevitably suffer from inherent texture biases, entangling essential structural details with high-frequency background noise. Furthermore, lacking cross-view geometric constraints, single-view metrics tend to overfit this noise, resulting in structural instability under few-shot conditions. To address these issues, this paper proposes the Frequency-Enhanced Dual-Subspace Network (FEDSNet). Specifically, FEDSNet utilizes the Discrete Cosine Transform (DCT) and a low-pass filtering mechanism to explicitly isolate low-frequency global structural components from spatial features, thereby suppressing background interference. Truncated Singular Value Decomposition (SVD) is employed to construct independent, low-rank linear subspaces for both spatial texture and frequency structural features. An adaptive gating mechanism is designed to dynamically fuse the projection distances from these dual views. This strategy leverages the structural stability of the frequency subspace to prevent the spatial subspace from overfitting to background features. Extensive experiments on four benchmark datasets - CUB-200-2011, Stanford Cars, Stanford Dogs, and FGVC-Aircraft - demonstrate that FEDSNet exhibits excellent classification performance and robustness, achieving highly competitive results compared to existing metric learning algorithms. Complexity analysis further confirms that the proposed network achieves a favorable balance between high accuracy and computational efficiency, providing an effective new paradigm for few-shot fine-grained visual recognition.

细粒度识别少样本学习频域分析度量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。