arXiv:2603.11783cs.CVcs.AI2026-03中稿 · and presented at R…

提出新框架提升遥感图像多标签分类,尤其在标签少时表现更好。

HELM: Hierarchical and Explicit Label Modeling with Graph Learning for Multi-Label Image Classification

  • 用视觉变换器中的层级专属类别标记捕捉标签间复杂关系。
  • 通过图卷积网络显式建模层级结构,生成具备层级感知的特征表示。
  • 引入自监督分支,有效利用无标签遥感图像数据,适合小样本场景。

层级多标签分类(HMLC)在遥感图像中对复杂标签依赖关系建模至关重要。现有方法在多路径层级结构下表现不佳,且很少利用无标签数据。本文提出HELM(Hierarchical and Explicit Label Modeling)框架,克服上述问题:(i) 在视觉变换器中使用层级专属类别标记,以捕捉精细的标签交互;(ii) 采用图卷积网络显式编码层级结构,生成具有层级感知的嵌入;(iii) 集成自监督分支,高效利用无标签遥感图像。在四个遥感图像数据集(UCM、AID、DFC-15、MLRSNet)上进行综合评估,HELM在监督与半监督设置下均达到最优性能,尤其在低标签场景中表现突出。

原文摘要 · Abstract (English)

Hierarchical multi-label classification (HMLC) is essential for modeling complex label dependencies in remote sensing. Existing methods, however, struggle with multi-path hierarchies where instances belong to multiple branches, and they rarely exploit unlabeled data. We introduce HELM (\textit{Hierarchical and Explicit Label Modeling}), a novel framework that overcomes these limitations. HELM: (i) uses hierarchy-specific class tokens within a Vision Transformer to capture nuanced label interactions; (ii) employs graph convolutional networks to explicitly encode the hierarchical structure and generate hierarchy-aware embeddings; and (iii) integrates a self-supervised branch to effectively leverage unlabeled imagery. We perform a comprehensive evaluation on four remote sensing image (RSI) datasets (UCM, AID, DFC-15, MLRSNet). HELM achieves state-of-the-art performance, consistently outperforming strong baselines in both supervised and semi-supervised settings, demonstrating particular strength in low-label scenarios.

多标签分类遥感图像自监督图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。