不需训练,用得分融合提升视觉语言模型零样本多标签识别性能
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
- 基于物体共现设计复合提示,利用大模型先验知识生成合理组合
- 发现最高得分反而不如次高分,提出去偏与分数融合算法修正偏差
- 可适配现有零样本方法,适合无标注数据场景的多标签识别任务
零样本多标签识别(MLR)在无训练数据、无需模型调优或结构修改的情况下面临挑战。现有方法依赖提示调优或架构调整,限制了零样本适用性。本文提出一种新方法,将视觉语言模型视为黑箱,仅利用输出得分而无需训练数据或真实标签。借助大语言模型对物体共现的先验知识,构建基于真实物体组合的复合提示。分析提示得分发现模型存在偏差及‘与’‘或’信号模糊问题,令人惊讶的是,最高复合得分远不如次高分。为此,我们设计去偏与得分融合算法,纠正图像偏差并澄清模型响应行为。该方法可增强其他零样本方法,持续提升性能。实验表明,其均值平均精度(mAP)优于需训练数据的方法,通过优化物体排序实现鲁棒零样本多标签识别。
原文摘要 · Abstract (English)
Zero-shot multi-label recognition (MLR) with Vision-Language Models (VLMs) faces significant challenges without training data, model tuning, or architectural modifications. Existing approaches require prompt tuning or architectural adaptations, limiting zero-shot applicability. Our work proposes a novel solution treating VLMs as black boxes, leveraging scores without training data or ground truth. Using large language model insights on object co-occurrence, we introduce compound prompts grounded in realistic object combinations. Analysis of these prompt scores reveals VLM biases and ``AND''/``OR'' signal ambiguities, notably that maximum compound scores are surprisingly suboptimal compared to second-highest scores. We address these through a debiasing and score-fusion algorithm that corrects image bias and clarifies VLM response behaviors. Our method enhances other zero-shot approaches, consistently improving their results. Experiments show superior mean Average Precision (mAP) compared to methods requiring training data, achieved through refined object ranking for robust zero-shot MLR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。