通过最大化特征维度熵,提升视觉语言模型在测试时提示调优的校准性。
D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
- 针对跨模态主导维度导致的校准偏差,提出维度熵最大化正则化方法。
- 在多个数据集上显著降低测试时提示调优的校准误差,提升预测可靠性。
- 适用于需要高可信度输出的现实部署场景,如医疗或自动驾驶视觉系统。
测试时自适应范式通过在未标注的目标数据上即时调整源模型,提升了对域偏移的灵活性。视觉语言模型(VLMs)凭借其泛化能力,可处理多样下游任务,而测试时提示调优已成为适配VLMs的主流方案。本文研究对比型VLMs,发现跨模态存在单一主导特征维度,该维度在文本与图像中均表现出高预测敏感性。约束其影响有助于改善校准误差。基于此,我们提出维度熵最大化方法,通过将文本特征分布推向均匀性,缓解主导维度的依赖问题。该方法有效缓解了测试时提示调优中的校准性能退化,为提升VLMs在真实场景部署中的可靠性提供了简单而有效的解决方案。
原文摘要 · Abstract (English)
Test-time adaptation paradigm provides flexibility towards domain shifts by performing immediate adaptation on unlabeled target data from the source model. Vision-Language Models (VLMs) leverage their generalization capabilities for diverse downstream tasks, and test-time prompt tuning has emerged as a prominent solution for adapting VLMs. In this work, we explore contrastive VLMs and identify the modality gap caused by a single dominant feature dimension across modalities. We observe that the dominant dimensions in both text and image modalities exhibit high predictive sensitivity, and that constraining their influence can improve calibration error. Building on this insight, we propose dimensional entropy maximization that regularizes the distribution of textual features toward uniformity to mitigate the dependency of dominant dimensions. Our method alleviates the degradation of calibration performance in test-time prompt tuning, offering a simple yet effective solution to enhance the reliability of VLMs in real-world deployment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。