用语言嵌入做信息瓶颈,让模型忽略环境干扰,提升跨域泛化能力。
Domain Generalization via Text-Anchored Information Bottleneck

- 以语言嵌入为锚点构建信息瓶颈,抑制视觉中的环境特异性特征。
- 在多个骨干网络上实现当前最优的域泛化性能,优于依赖视觉表达的方法。
- 适合关注模型鲁棒性、想摆脱环境偏差的研究者或工程落地场景。
视觉识别模型在新环境中常失效。域泛化(DG)通过学习对环境变化不变的表示来应对这一问题。现有方法越来越多地依赖大视觉语言模型,假设保留其丰富的视觉表示能提升鲁棒性。然而我们发现,这种视觉表达反而会传播与训练环境相关的虚假线索,阻碍不变性学习。因此,我们放弃视觉引导,转而将语言嵌入空间视为域不变性的主要来源,自然形成信息瓶颈,保留核心语义并抑制域特异性差异。在多种骨干网络上的大量实验表明,该方法达到当前最优性能,并进一步分析了有效指导的关键因素。这些发现将DG重点从改进表示转向设计强制不变性的监督机制。
原文摘要 · Abstract (English)
Visual recognition models often fail when deployed in new environments. Domain Generalization (DG) addresses this by learning representations that remain invariant to environment-specific variations. Recent approaches increasingly rely on large vision-language models, assuming that preserving their expressive visual representations improves robustness. However, we show that such visual expressiveness can instead propagate spurious cues that tie representations to the training environments, hindering invariant learning. We therefore discard visual guidance and instead treat the language embedding space as the primary source of domain invariance, naturally acting as an information bottleneck that preserves core semantics while suppressing domain-specific variations. Extensive experiments across diverse backbones exhibit state-of-the-art performance and further analyze what makes guidance effective for robust generalization. These findings shift the focus of DG from improving representations to designing supervision that enforces invariance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。