提出新方法提升视觉语言模型测试时提示调优的置信度校准能力。
SoC: Semantic Orthogonal Calibration for Test-Time Prompt Tuning
- 引入基于Huber损失的语义正交校准,平衡类别分离与语义相近性。
- 在多个数据集上显著改善置信度校准,同时保持分类性能不下降。
- 适合关注模型可靠性、需精准置信度估计的医疗、自动驾驶等场景。
随着视觉语言模型(VLMs)在医疗、自动驾驶等关键决策系统中的广泛应用,其不确定性估计的校准变得至关重要。然而,现有测试时提示调优(TPT)研究主要聚焦于提升判别性能,忽视了校准问题。最新方法主张对文本提示嵌入强制完全正交以增强可分性,从而提升校准效果。但本文理论分析表明,完全正交约束的梯度会强烈推动语义相关类别彼此远离,导致模型过度自信。为此,我们提出语义正交校准(SoC),一种基于Huber损失的正则化方法,实现平滑原型分离的同时保留语义邻近性,显著优于以往基于正交性的方法。在全面实证验证中,SoC consistently 提升校准性能,并保持优异的判别能力。
原文摘要 · Abstract (English)
With the increasing adoption of vision-language models (VLMs) in critical decision-making systems such as healthcare or autonomous driving, the calibration of their uncertainty estimates becomes paramount. Yet, this dimension has been largely underexplored in the VLM test-time prompt-tuning (TPT) literature, which has predominantly focused on improving their discriminative performance. Recent state-of-the-art advocates for enforcing full orthogonality over pairs of text prompt embeddings to enhance separability, and therefore calibration. Nevertheless, as we theoretically show in this work, the inherent gradients from fully orthogonal constraints will strongly push semantically related classes away, ultimately making the model overconfident. Based on our findings, we propose Semantic Orthogonal Calibration (SoC), a Huber-based regularizer that enforces smooth prototype separation while preserving semantic proximity, thereby improving calibration compared to prior orthogonality-based approaches. Across a comprehensive empirical validation, we demonstrate that SoC consistently improves calibration performance, while also maintaining competitive discriminative capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。