提出新方法缓解视觉语言模型测试时适应中的偏见问题
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
- 通过解耦增强探索与公平校准,避免依赖熵最小化
- 在多个领域迁移和细粒度分类任务上表现优于现有方法
- 适合关注模型鲁棒性与公平性的研究人员
视觉语言模型(如CLIP)虽具备强大的零样本识别能力,但在分布偏移下性能显著下降。测试时自适应(TTA)旨在仅使用无标签测试样本提升鲁棒性,但多数基于提示的TTA方法依赖熵最小化,当类别共享视觉特征时可能放大虚假相关性并导致过度自信错误。本文提出公平上下文学习(FCL),一种基于事件的TTA框架,通过显式处理共享证据偏差避免熵最小化。受可加证据分解假设启发,FCL将适配过程分为两步:(i) 增强驱动探索以识别合理类别候选;(ii) 公平性驱动校准,使文本上下文对共有的视觉证据敏感度均等。该公平约束缓解了对部分特征的过度关注,实现无需依赖熵减的文本嵌入有效校准。通过大量实验验证,FCL在多种领域迁移与细粒度基准上达到与先进TTA方法相当的性能。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) such as CLIP enable strong zero-shot recognition but suffer substantial degradation under distribution shifts. Test-Time Adaptation (TTA) aims to improve robustness using only unlabeled test samples, yet most prompt-based TTA methods rely on entropy minimization -- an approach that can amplify spurious correlations and induce overconfident errors when classes share visual features. We propose Fair Context Learning (FCL), an episodic TTA framework that avoids entropy minimization by explicitly addressing shared-evidence bias. Motivated by our additive evidence decomposition assumption, FCL decouples adaptation into (i) augmentation-based exploration to identify plausible class candidates, and (ii) fairness-driven calibration that adapts text contexts to equalize sensitivity to common visual evidence. This fairness constraint mitigates partial feature obsession and enables effective calibration of text embeddings without relying on entropy reduction. Through extensive evaluation, we empirically validate our theoretical motivation and show that FCL achieves competitive adaptation performance relative to state-of-the-art TTA methods across diverse domain-shift and fine-grained benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。