揭秘CLIP测试时自适应的真正驱动力,发现轻量更新更有效。
What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective

- 按测试时更新内容将方法分为三类,统一分析框架。
- 适应增益主要来自测试证据和可靠代理,而非复杂优化。
- 不同数据偏移下最优策略不同,无万能方案。
视觉-语言模型(如CLIP)已成为开放词汇识别的标准骨干,但其零样本预测在部署时仍易受分布偏移影响。测试时自适应(TTA)作为轻量级解决方案被引入到CLIP中,催生了大量TTA4CLIP方法。然而,该领域的实证进展远超我们对其真正驱动因素、性能提升来源及可靠性边界的理解。本文回归基础,对TTA4CLIP开展系统性控制实验研究。首先,根据测试时更新内容将现有方法归纳为三类统一范式;随后提出TTABC——一个开源的CLIP TTA基准,标准化评估协议并集成20余种代表性方法。我们的控制实验聚焦三个关键方面:第一,揭示参数类方法的驱动因素,发现适应增益主要源自测试时证据与可靠代理,而非高强度优化;第二,探索超越复杂参数调优的证据利用方式,表明通过跨样本或当前样本证据及轻量原型更新即可实现高效竞争性能;第三,证明不存在通用最优方案:单一适配范式无法适用于所有偏移场景,最佳策略取决于偏移类型。我们希望本研究与基准能为理解当前TTA4CLIP格局提供清晰认知,并为后续研究奠定基础。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts encountered at deployment. Test-Time Adaptation (TTA) has recently been extended to CLIP as a lightweight solution, leading to a rapidly growing body of TTA4CLIP methods. However, empirical progress in this area has largely outpaced our understanding of what truly drives adaptation, where their gains originate, and under which shifts they remain reliable. In this paper, we take a step back from the pursuit of state-of-the-art accuracy and conduct a systematic controlled study of TTA4CLIP. We first organize existing methods into three unified paradigms according to what is updated at test time. We then introduce TTABC, an open-source TTA Benchmark for CLIP, which standardizes evaluation protocols and integrates more than 20 representative methods. Our controlled empirical analysis focuses on three key areas. First, we determine the driving factors in parameter-based methods, revealing that adaptation gains are primarily driven by test-time evidence and reliable proxies rather than heavy optimization. Second, we explore evidence utilization beyond heavy parameter tuning, showing that competitive and efficient performance can be achieved through cross- or current-sample evidence and lightweight prototype updates. Finally, we demonstrate that there is no silver bullet for TTA: no single adaptation paradigm is universally optimal, and the preferred paradigm depends on the nature of shift. We hope our benchmark and study provide a clearer understanding of the current TTA4CLIP landscape and establish a foundation for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。