arXiv:2412.19491cs.CV2024-12被引 2

通过多阶上下文信息提升图像多标签分类精度

Multi-label Classification using Deep Multi-order Context-aware Kernel Networks

  • 设计可学习的上下文感知核函数,融合图像块间多距离邻域关系
  • 在Corel5K和NUS-WIDE上达到与顶尖方法相当的准确率
  • 适合关注视觉上下文建模的多标签图像分类研究者

多标签分类是模式识别中的挑战性任务。尽管深度学习方法已显著提升性能,但多数先进方法忽略了模型学习过程中的上下文信息。由于上下文可能为模型提供额外线索,显著提升分类效果。本文充分利用图像的几何结构等上下文信息,学习更优的上下文感知相似性(即核函数)。将上下文感知核设计重构为前馈网络,输出显式核映射特征。所提出的深度多阶上下文感知核网络(DMCKN)进一步利用不同距离下的多阶图像块邻域,实现更优的判别能力。在Corel5K和NUS-WIDE两个基准数据集上评估,实验结果表明该方法在定量与定性层面均表现优异,优于现有主流方法。

原文摘要 · Abstract (English)

Multi-label classification is a challenging task in pattern recognition. Many deep learning methods have been proposed and largely enhanced classification performance. However, most of the existing sophisticated methods ignore context in the models' learning process. Since context may provide additional cues to the learned models, it may significantly boost classification performances. In this work, we make full use of context information (namely geometrical structure of images) in order to learn better context-aware similarities (a.k.a. kernels) between images. We reformulate context-aware kernel design as a feed-forward network that outputs explicit kernel mapping features. Our obtained context-aware kernel network further leverages multiple orders of patch neighbors within different distances, resulting into a more discriminating Deep Multi-order Context-aware Kernel Network (DMCKN) for multi-label classification. We evaluate the proposed method on the challenging Corel5K and NUS-WIDE benchmarks, and empirical results show that our method obtains competitive performances against the related state-of-the-art, and both quantitative and qualitative performances corroborate its effectiveness and superiority for multi-label image classification.

多标签分类上下文建模深度学习图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。