arXiv:2502.19946cs.CV2025-02被引 1

无需训练,通过基变换旋转特征空间提升视觉语言模型测试时适应能力

Space Rotation with Basis Transformation for Training-free Test-Time Adaptation

  • 用基变换重构特征空间,增强类别间差异
  • 动态队列存储代表性样本,提升信息捕捉能力
  • 完全免训练,兼顾性能与效率,适合部署场景

随着视觉-语言模型(VLM)在下游任务中的应用发展,基于VLM的测试时自适应方法因其能应对测试阶段分布变化而受到关注。尽管已有方法取得一定进展,但通常需要大量计算资源或受限于原始特征空间,难以有效支持测试时自适应。为此,本文提出一种免训练的特征空间旋转方法,通过基变换重构原始特征空间,映射到新表示,从而增强类别间区分度,为测试阶段提供更有效指导。同时,设计动态队列以存储各类别代表性样本,更好捕获相关特征。在多个基准上的实验表明,该方法在性能和效率上均优于现有最先进技术。

原文摘要 · Abstract (English)

With the development of visual-language models (VLM) in downstream task applications, test-time adaptation methods based on VLM have attracted increasing attention for their ability to address changes distribution in test-time. Although prior approaches have achieved some progress, they typically either demand substantial computational resources or are constrained by the limitations of the original feature space, rendering them less effective for test-time adaptation tasks. To address these challenges, we propose a training-free feature space rotation with basis transformation for test-time adaptation. By leveraging the inherent distinctions among classes, we reconstruct the original feature space and map it to a new representation, thereby enhancing the clarity of class differences and providing more effective guidance for the model during testing. Additionally, to better capture relevant information from various classes, we maintain a dynamic queue to store representative samples. Experimental results across multiple benchmarks demonstrate that our method outperforms state-of-the-art techniques in terms of both performance and efficiency.

测试时适应视觉语言模型特征空间变换免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。