arXiv:2503.12096cs.CV2025-03CVPR被引 19

通过正交约束提升视觉语言模型测试时提示调优的可靠性

O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models

  • 在可学习提示的文本特征上施加正交约束,改善校准性能
  • 在多个数据集上显著降低平均校准误差,优于现有最优方法
  • 特别适合对模型置信度敏感的细粒度分类任务

视觉语言模型(VLMs)的测试时提示调优因其无需微调即可利用无标签数据而受到关注。尽管该方法能提升准确率,但模型校准性能较差,影响其可靠性。本文提出O-TPT方法,在可学习提示对应的文本特征上引入正交性约束,以改进测试时提示调优的校准效果。我们揭示了现有依赖文本特征离散度的方法存在局限,并证明简单的正交化处理能更有效地实现文本特征离散。在多种骨干网络和基线设置下进行的大量实验表明,本方法在降低整体平均校准误差方面持续优于现有最先进方法,且在细粒度分类任务中超越零样本校准表现。

原文摘要 · Abstract (English)

Test-time prompt tuning for vision-language models (VLMs) is getting attention because of their ability to learn with unlabeled data without fine-tuning. Although test-time prompt tuning methods for VLMs can boost accuracy, the resulting models tend to demonstrate poor calibration, which casts doubts on the reliability and trustworthiness of these models. Notably, more attention needs to be devoted to calibrating the test-time prompt tuning in vision-language models. To this end, we propose a new approach, called O-TPT that introduces orthogonality constraints on the textual features corresponding to the learnable prompts for calibrating test-time prompt tuning in VLMs. Towards introducing orthogonality constraints, we make the following contributions. First, we uncover new insights behind the suboptimal calibration performance of existing methods relying on textual feature dispersion. Second, we show that imposing a simple orthogonalization of textual features is a more effective approach towards obtaining textual dispersion. We conduct extensive experiments on various datasets with different backbones and baselines. The results indicate that our method consistently outperforms the prior state of the art in significantly reducing the overall average calibration error. Also, our method surpasses the zero-shot calibration performance on fine-grained classification tasks.

提示调优模型校准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。