提出新基准评估视觉语言模型测试时自适应,发现现有方法提升有限且损害模型可信度
The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models

- 构建统一框架评测8种实时适应与7种在线适应方法
- 发现多数方法增益微弱,且与训练调优方法不兼容
- 揭示准确率提升常伴随可信度下降,适合关注模型可靠性研究者
测试时自适应(TTA)方法近年来受到广泛关注,可在推理阶段提升视觉语言模型(如CLIP)性能,无需额外标注数据。然而,当前研究普遍存在基线结果重复、评估指标单一、实验设置不一致和分析不足等问题,阻碍了不同TTA方法间的公平比较,也难以评估其实际优劣。为此,我们提出TTA-VLM,一个针对视觉语言模型的综合性基准。该基准在统一可复现框架中实现了8种实时适应和7种在线适应方法,并在15个常用数据集上进行评估。不同于以往仅聚焦CLIP的研究,我们扩展至使用Sigmoid损失训练的SigLIP模型,并引入CoOp、MaPLe、TeCoA等训练时调优方法以评估泛化性。除分类准确率外,还引入鲁棒性、校准度、分布外检测和稳定性等多维度指标,实现更全面的评估。大量实验表明:1)现有TTA方法相较于早期工作提升有限;2)当前方法与训练时调优方法协同效果差;3)准确率提升常以模型可信度降低为代价。我们开源TTA-VLM,旨在促进更公平、全面的评估,并推动社区发展更可靠、通用的TTA策略。
原文摘要 · Abstract (English)
Test-time adaptation (TTA) methods have gained significant attention for enhancing the performance of vision-language models (VLMs) such as CLIP during inference, without requiring additional labeled data. However, current TTA researches generally suffer from major limitations such as duplication of baseline results, limited evaluation metrics, inconsistent experimental settings, and insufficient analysis. These problems hinder fair comparisons between TTA methods and make it difficult to assess their practical strengths and weaknesses. To address these challenges, we introduce TTA-VLM, a comprehensive benchmark for evaluating TTA methods on VLMs. Our benchmark implements 8 episodic TTA and 7 online TTA methods within a unified and reproducible framework, and evaluates them across 15 widely used datasets. Unlike prior studies focused solely on CLIP, we extend the evaluation to SigLIP--a model trained with a Sigmoid loss--and include training-time tuning methods such as CoOp, MaPLe, and TeCoA to assess generality. Beyond classification accuracy, TTA-VLM incorporates various evaluation metrics, including robustness, calibration, out-of-distribution detection, and stability, enabling a more holistic assessment of TTA methods. Through extensive experiments, we find that 1) existing TTA methods produce limited gains compared to the previous pioneering work; 2) current TTA methods exhibit poor collaboration with training-time fine-tuning methods; 3) accuracy gains frequently come at the cost of reduced model trustworthiness. We release TTA-VLM to provide fair comparison and comprehensive evaluation of TTA methods for VLMs, and we hope it encourages the community to develop more reliable and generalizable TTA strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。