arXiv:2604.19339cs.CV2026-04

通过分解整体特征提升小样本细粒度识别能力

Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data

论文配图:Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data
图 1 · 摘自论文原文
  • 将整体特征拆解为局部细微差异,逐步构建整体认知
  • 在5个数据集上显著优于现有方法,提升识别精度
  • 适合小样本、高相似度图像分类任务研究者

超细粒度视觉分类(Ultra-FGVC)旨在用少量训练样本对高度相似的子类别进行分类。然而,如在极相似品种中叶片轮廓这类整体判别性线索,仍未被充分探索,制约了识别性能。尽管关键,建模具有复杂形态结构的整体线索通常需大量数据,在数据受限场景下面临挑战。为此,本文提出新型分治整体认知网络(DHCNet),通过将整体线索分解为空间关联的细微差异,并逐步建立整体认知过程,显著简化整体认知并降低对训练数据的依赖。技术上,DHCNet从局部小区域开始,利用自打乱操作逐步分析更大型区域的细微差异;同时借助未受影响的局部区域,潜在引导对打乱块间原始拓扑结构的感知,辅助建立这些差异的空间关联。此外,网络将从局部区域发现的整体线索在线优化并融入训练,迭代提升其质量。最终,以这些整体线索作为监督信号,微调识别模型参数,增强其对整体线索的敏感性。大量实验表明,DHCNet在五个广泛使用的Ultra-FGVC数据集上表现优异。

原文摘要 · Abstract (English)

Ultra-fine-grained visual categorization (Ultra-FGVC) aims to classify highly similar subcategories within fine-grained objects using limited training samples. However, holistic yet discriminative cues, such as leaf contours in extremely similar cultivars, remain under-explored in current studies, thereby limiting recognition performance. Though crucial, modeling holistic cues with complex morphological structures typically requires massive training samples, posing significant challenges in data-limited scenarios. To address this challenge, we propose a novel Divide-and-Conquer Holistic Cognition Network (DHCNet) that implements a divide-and-conquer strategy by decomposing holistic cues into spatially-associated subtle discrepancies and progressively establishing the holistic cognition process, significantly simplifying holistic cognition while reducing dependency on training data. Technically, DHCNet begins by progressively analyzing subtle discrepancies, transitioning from smaller local patches to larger ones using a self-shuffling operation on local regions. Simultaneously, it leverages the unaffected local regions to potentially guide the perception of the original topological structure among the shuffled patches, thereby aiding in the establishment of spatial associations for these discrepancies. Additionally, DHCNet incorporates the online refinement of these holistic cues discovered from local regions into the training process to iteratively improve their quality. As a result, DHCNet uses these holistic cues as supervisory signals to fine-tune the parameters of the recognition model, thus improving its sensitivity to holistic cues across the entire objects. Extensive evaluations demonstrate that DHCNet achieves remarkable performance on five widely-used Ultra-FGVC datasets.

细粒度分类小样本学习整体认知图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。