arXiv:2608.22313cs.CV2026-08TPAMI

用语言引导的自适应机制,提升部分标注下的多标签图像分类性能。

Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label

论文配图:Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label
图 1 · 摘自论文原文
  • 基于CLIP构建密集视觉对比约束,迁移任务知识到视觉域。
  • 通过类特定提示调优实现语言与视觉域交互适配,提升语义恢复精度。
  • 在公开数据集上达新最优,且可解释性强,适合弱监督学习研究者。

在不完整标注下进行多标签图像分类是一项具有挑战性的任务,因其在大规模数据集上兼具高效率与低人工成本的优势而受到广泛关注。现有主流方法依赖强先验假设从部分标注中恢复缺失语义,但这些统计先验易导致语义错误,引发灾难性过拟合。为此,本文提出语言驱动的密集语义适配器(LDSA),从多模态预训练模型CLIP中挖掘先验自适应关系。首先提出密集对比适配器,构建密集视觉对比约束,将任务知识迁移至视觉域;随后引入语言驱动的交互解码器,结合类特定提示调优,使语言代理与视觉域自适应对齐。在协同学习机制下,实验表明所提LDSA在公开多标签分类基准上达到新最优性能,可解释性分析揭示其通过先验自适应学习发现了隐含语义关系。

原文摘要 · Abstract (English)

Learning multi-label image classification with incomplete annotations is a challenging task that has been widely studied for its superior trade-off between high efficiency and less labor consumption on large-scale datasets. Predominant methods rely on strong prior assumptions to recover the missing semantics from partial annotations. However, these statistic priors suffer from unstable semantic mistakes and thus lead to catastrophic overfitting. Toward this end, we propose a Language-driven Dense Semantic Adaptor (LDSA) that excavates prior-adaptive relationships from multimodal pretrained CLIP models. In our approach, the densely contrastive adaptor is first proposed to construct dense visual contrastive constraints, transferring the task-specific knowledge to visual domains. We then propose a language-driven interactive decoder with the help of class-specific prompt tuning, which adapts language proxies with visual domains. With the collaborative learning of proposed modules, experimental results demonstrate our proposed LDSA achieves a new state of the art on public multi-label classification benchmarks, and interpretable analyses reveal that our LDSA discovers implicit semantic relationships with the prior-adaptive learning scheme.

多标签分类弱监督学习视觉语言模型提示调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。