arXiv:2512.18411cs.CVcs.AI2025-12IJCV被引 1

提出新方法缓解视觉语言模型中的提示偏差,提升少样本适应能力。

AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning

论文配图:AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
图 1 · 摘自论文原文
  • 通过自适应去偏的集成学习融合多提示预测结果
  • 在三个任务上均超越现有方法,显著提升泛化性能
  • 适合需要快速适配新任务的视觉语言模型研究者

多提示学习已成为高效适应视觉语言模型至下游任务的有效方法,尤其在资源有限时。现有方法多聚焦于单一基础模型内设计精巧提示以提升性能,却忽略了模型-提示匹配偏差:同一提示在不同模型(如CLIP-ViT-B/16与CLIP-ViT-B/32)中语义不同,导致相同提示产生不一致预测。为此,本文提出集成学习策略,充分聚合多样预测优势。同时揭示了样本-提示匹配偏差的存在——输入样本中包含与提示无关的语义信息,直接利用全部样本信息生成集成权重会降低性能。因此,基于信息论分析,提取样本中与提示相关的信息,自适应计算去偏的集成权重。整体提出自适应去偏集成多提示学习(AmPLe),同时缓解两类偏差。在三个代表性任务(新类别泛化、新目标数据集、未见域偏移)上的大量实验表明,AmPLe广泛优于现有方法。因果视角的理论验证进一步支持其有效性。

原文摘要 · Abstract (English)

Multi-prompt learning methods have emerged as an effective approach for facilitating the rapid adaptation of vision-language models to downstream tasks with limited resources. Existing multi-prompt learning methods primarily focus on utilizing various meticulously designed prompts within a single foundation vision-language model to achieve superior performance. However, the overlooked model-prompt matching bias hinders the development of multi-prompt learning, i.e., the same prompt can convey different semantics across distinct vision-language models, such as CLIP-ViT-B/16 and CLIP-ViT-B/32, resulting in inconsistent predictions of identical prompt. To mitigate the impact of this bias on downstream tasks, we explore an ensemble learning approach to sufficiently aggregate the benefits of diverse predictions. Additionally, we further disclose the presence of sample-prompt matching bias, which originates from the prompt-irrelevant semantics encapsulated in the input samples. Thus, directly utilizing all information from the input samples for generating weights of ensemble learning can lead to suboptimal performance. In response, we extract prompt-relevant semantics from input samples by leveraging the guidance of the information theory-based analysis, adaptively calculating debiased ensemble weights. Overall, we propose Adaptive-Debiased Ensemble MultiPrompt Learning, abbreviated as AmPLe, to mitigate the two types of bias simultaneously. Extensive experiments on three representative tasks, i.e., generalization to novel classes, new target datasets, and unseen domain shifts, show that AmPLe can widely outperform existing methods. Theoretical validation from a causal perspective further supports the effectiveness of AmPLe.

视觉语言模型多提示学习去偏方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。