arXiv:2511.10260cs.CVcs.AI2025-11中稿 · IEEE Transactions …

用超图建模细粒度视觉分类中的语义关系,提升特征聚合效果。

H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification

  • 通过加权超图动态聚合局部特征,捕捉高阶语义依赖。
  • 在四个基准数据集上达到新纪录,显著提升分类精度。
  • 适合细粒度图像识别、视觉模型设计等研究者参考。

细粒度视觉分类(FGVC)因类别间差异微小、类内变化大而极具挑战。现有方法多依赖特征选择或区域提议来定位判别性区域,但常无法全面捕捉判别线索,并引入大量与类别无关的冗余信息。为此,我们提出H3Former,一种基于超图的标记到区域框架,利用高阶语义关系对局部细粒度表示进行结构化区域级建模。具体地,我们设计了语义感知聚合模块(SAAM),通过多尺度上下文线索动态构建标记间的加权超图,并利用超图卷积捕获高阶语义依赖,逐步将标记特征聚合为紧凑的区域级表示。此外,我们提出双曲分层对比损失(HHCL),在非欧氏嵌入空间中施加分层语义约束,增强类间可分性和类内一致性,同时保持细粒度类别间的内在层次关系。在四个标准FGVC基准上的实验验证了所提H3Former框架的优越性。

原文摘要 · Abstract (English)

Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on feature-selection mechanisms or region-proposal strategies to localize discriminative regions for semantic analysis. However, these methods often fail to capture discriminative cues comprehensively while introducing substantial category-agnostic redundancy. To address these limitations, we propose H3Former, a novel token-to-region framework that leverages high-order semantic relations to aggregate local fine-grained representations with structured region-level modeling. Specifically, we propose the Semantic-Aware Aggregation Module (SAAM), which exploits multi-scale contextual cues to dynamically construct a weighted hypergraph among tokens. By applying hypergraph convolution, SAAM captures high-order semantic dependencies and progressively aggregates token features into compact region-level representations. Furthermore, we introduce the Hyperbolic Hierarchical Contrastive Loss (HHCL), which enforces hierarchical semantic constraints in a non-Euclidean embedding space. The HHCL enhances inter-class separability and intra-class consistency while preserving the intrinsic hierarchical relationships among fine-grained categories. Comprehensive experiments conducted on four standard FGVC benchmarks validate the superiority of our H3Former framework.

细粒度分类超图建模语义聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。