无需训练,实时调整视觉语言模型预测,提升跨域识别准确率。
Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM
- 用在线EM算法结合零样本预测,动态更新测试样本后验概率。
- 在15个数据集上显著优于现有方法,跨域和分布外场景均稳定提升。
- 不依赖历史数据或标注,适合实际部署中的快速适应场景。
视觉-语言模型(VLMs)在开放世界图像识别中表现出强大泛化能力,但其实际应用受领域偏移和分布变化影响。为此,测试时自适应(TTA)成为主流范式,利用测试时的现成数据进行独立样本预测,无需测试标注。然而,传统方法常需代价高昂的训练或优化,或假设可访问/存储历史数据。本文提出FreeTTA,一种无需训练、通用且无假设的TTA方法,首次显式建模测试数据分布,利用测试样本间的内在关系增强单个样本预测,而无需同时访问。FreeTTA通过在线EM算法,以VLM的零样本预测作为先验,迭代计算每个在线测试样本的后验概率并更新参数。实验表明,该方法在15个数据集上跨域与分布外设置下均显著优于当前最优方法。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have become prominent in open-world image recognition for their strong generalization abilities. Yet, their effectiveness in practical applications is compromised by domain shifts and distributional changes, especially when test data distributions diverge from training data. Therefore, the paradigm of test-time adaptation (TTA) has emerged, enabling the use of online off-the-shelf data at test time, supporting independent sample predictions, and eliminating reliance on test annotations. Traditional TTA methods, however, often rely on costly training or optimization processes, or make unrealistic assumptions about accessing or storing historical training and test data. Instead, this study proposes FreeTTA, a training-free and universally available method that makes no assumptions, to enhance the flexibility of TTA. More importantly, FreeTTA is the first to explicitly model the test data distribution, enabling the use of intrinsic relationships among test samples to enhance predictions of individual samples without simultaneous access--a direction not previously explored. FreeTTA achieves these advantages by introducing an online EM algorithm that utilizes zero-shot predictions from VLMs as priors to iteratively compute the posterior probabilities of each online test sample and update parameters. Experiments demonstrate that FreeTTA achieves stable and significant improvements compared to state-of-the-art methods across 15 datasets in both cross-domain and out-of-distribution settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。