arXiv:2509.17747cs.CVcs.AI2025-09中稿 · IEEE Transactions …被引 5

用双视角对齐与分层提示缓解多标签图像分类中的类别不平衡问题。

Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification

  • 通过双视角对齐学习提取互补特征,增强图文对齐能力。
  • 在长尾和少样本场景下,mAP提升最高达10.0%和6.8%。
  • 适合处理数据不均衡的多标签图像识别任务,如医疗影像分析。

现实世界数据集常呈现多类别上的类别不平衡,表现为长尾分布和少样本场景。这在类别不平衡多标签图像分类(CI-MLIC)任务中尤为严峻,因数据不平衡与多目标识别带来显著挑战。为此,我们提出一种新方法——基于分层提示的双视角对齐学习(HP-DVAL),利用视觉语言预训练(VLP)模型的多模态知识缓解多标签场景下的类别不平衡问题。具体而言,HP-DVAL采用双视角对齐学习,通过提取互补特征,将VLP模型的强大特征表示能力迁移至任务中。为更好适配VLP模型于CI-MLIC任务,引入分层提示调优策略,利用全局与局部提示学习任务特异性及上下文相关先验知识。此外,在提示调优阶段设计语义一致性损失,防止提示偏离嵌入于VLP模型中的通用知识。方法在两个CI-MLIC基准数据集MS-COCO与VOC2007上验证,实验结果表明其优于现有最先进方法:在长尾多标签图像分类任务中实现mAP提升10.0%和5.2%;在多标签少样本图像分类任务中提升6.8%和2.9%。

原文摘要 · Abstract (English)

Real-world datasets often exhibit class imbalance across multiple categories, manifesting as long-tailed distributions and few-shot scenarios. This is especially challenging in Class-Imbalanced Multi-Label Image Classification (CI-MLIC) tasks, where data imbalance and multi-object recognition present significant obstacles. To address these challenges, we propose a novel method termed Dual-View Alignment Learning with Hierarchical Prompt (HP-DVAL), which leverages multi-modal knowledge from vision-language pretrained (VLP) models to mitigate the class-imbalance problem in multi-label settings. Specifically, HP-DVAL employs dual-view alignment learning to transfer the powerful feature representation capabilities from VLP models by extracting complementary features for accurate image-text alignment. To better adapt VLP models for CI-MLIC tasks, we introduce a hierarchical prompt-tuning strategy that utilizes global and local prompts to learn task-specific and context-related prior knowledge. Additionally, we design a semantic consistency loss during prompt tuning to prevent learned prompts from deviating from general knowledge embedded in VLP models. The effectiveness of our approach is validated on two CI-MLIC benchmarks: MS-COCO and VOC2007. Extensive experimental results demonstrate the superiority of our method over SOTA approaches, achieving mAP improvements of 10.0\% and 5.2\% on the long-tailed multi-label image classification task, and 6.8\% and 2.9\% on the multi-label few-shot image classification task.

多标签分类类别不平衡提示调优视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。