无需训练和反向传播,用统计方法动态适应视觉语言模型新场景
Training-Free Test-Time Adaptation with Brownian Distance Covariance in Vision-Language Models
- 基于布朗距离协方差,通过成对距离捕捉跨模态依赖关系
- 在多个数据集上实现顶尖泛化性能,计算开销显著降低
- 适合需快速部署、无训练资源的视觉语言应用
视觉语言模型在域偏移下性能下降,限制了实际应用。现有测试时自适应方法计算量大、依赖反向传播,且多聚焦单一模态。为此,我们提出无需训练的测试时自适应方法 TaTa,利用布朗距离协方差——一种通过成对距离捕捉线性和非线性依赖关系的强大统计量——在不进行训练或反向传播的情况下动态适配视觉语言模型至新域。该方法不仅提升效率,还通过避免权重剧烈更新增强了稳定性。TaTa 进一步融合属性增强提示,利用描述性视觉线索改进视觉-语言推理。结合动态聚类与伪标签精炼,有效重校准模型以应对新型视觉情境。跨多个数据集的实验表明,TaTa 显著降低计算成本的同时,在域泛化与跨数据集泛化任务中达到当前最优表现。
原文摘要 · Abstract (English)
Vision-language models suffer performance degradation under domain shift, limiting real-world applicability. Existing test-time adaptation methods are computationally intensive, rely on back-propagation, and often focus on single modalities. To address these issues, we propose Training-free Test-Time Adaptation with Brownian Distance Covariance (TaTa). TaTa leverages Brownian Distance Covariance-a powerful statistical measure that captures both linear and nonlinear dependencies via pairwise distances-to dynamically adapt VLMs to new domains without training or back-propagation. This not only improves efficiency but also enhances stability by avoiding disruptive weight updates. TaTa further integrates attribute-enhanced prompting to improve vision-language inference with descriptive visual cues. Combined with dynamic clustering and pseudo-label refinement, it effectively recalibrates the model for novel visual contexts. Experiments across diverse datasets show that TaTa significantly reduces computational cost while achieving state-of-the-art performance in domain and cross-dataset generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。