arXiv:2606.07646cs.CVcs.AI2026-06

通过显式建模样本级域特征,提升测试时自适应的鲁棒性。

DOME: Learning Transferable Domain Variables from Sparse Supervision for Test-Time Adaptation

论文配图:DOME: Learning Transferable Domain Variables from Sparse Supervision for Test-Time Adaptation
图 1 · 摘自论文原文
  • 用视觉语言预训练提取连续域表示,参数化为分布变量
  • 零样本下利用稀疏域库实现解耦监督,性能超越复杂方法
  • 适合需要稳定测试时适配的工业部署场景

测试时自适应(TTA)旨在仅使用无标签流数据将模型对齐到变化的测试域。现有方法通常隐式推断单一全局域分布,忽视真实域偏移的多维性和样本特异性,导致适应效果脆弱。本文提出DOME,一种有效的域编码器,以零样本方式显式建模每个样本的域特征。DOME利用视觉语言预训练提取密集连续表征,将域参数化为分布变量,并引入动量更新的稀疏域库实现解耦监督。通过将这些显式域线索注入下游模型,即使基础的熵最小化TTA策略,在ImageNet-C、ImageNet-R和ImageNet-Sketch上也达到领先性能,优于复杂方法。结果表明,鲁棒适应的关键不在于复杂的算法,而在于显式、结构化的域表示。

原文摘要 · Abstract (English)

Test-time adaptation (TTA) aims to align a model to shifting test domains using only unlabeled streaming data. Most existing methods implicitly infer a single global domain distribution, ignoring the multidimensional and sample-specific nature of real-world domain shifts, leading to fragile adaptation. We propose DOME, an effective domain encoder that explicitly models each sample's domain in a zero-shot manner. DOME leverages vision-language pretraining to extract dense, continuous representations, parameterizes domains as distributional variables, and introduces a momentum-updated sparse domain bank for disentangled supervision. By injecting these explicit domain cues into downstream models, even a basic entropy-minimization TTA strategy achieves state-of-the-art performance across ImageNet-C, ImageNet-R, and ImageNet-Sketch, outperforming complex TTA approaches. Our results demonstrate that robust adaptation stems not from intricate adaptation algorithms, but from explicit, structured domain representation.

测试时自适应域泛化视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。