通过流分解建模风格不变特征,提升视觉领域泛化能力
DGFamba: Learning Flow Factorized State Space for Visual Domain Generalization
- 用流分解映射风格增强与原始特征,构建风格不变表示
- 在隐空间对齐概率路径,实现跨风格内容分布一致
- 在多个数据集上达到当前最优,适合做视觉泛化研究
领域泛化旨在从源域学习可泛化至任意未见目标域的表征。视觉领域泛化的关键挑战在于风格差异导致的域间差距,而图像内容保持稳定。以VMamba为代表的选通状态空间模型展现出全局感受野,能有效表征内容信息。然而,如何利用选通状态空间提取域不变特性尚未被充分探索。本文提出一种新型流分解状态空间模型DG-Famba,用于视觉领域泛化。为保持域一致性,创新性地通过流分解映射风格增强与原始状态嵌入。在隐空间中,每种风格的状态嵌入由潜在概率路径定义。通过对齐这些概率路径,状态嵌入可在不同风格下表示相同的内容分布。在多种视觉领域泛化设置下的大量实验表明,该方法性能达到当前最优水平。
原文摘要 · Abstract (English)
Domain generalization aims to learn a representation from the source domain, which can be generalized to arbitrary unseen target domains. A fundamental challenge for visual domain generalization is the domain gap caused by the dramatic style variation whereas the image content is stable. The realm of selective state space, exemplified by VMamba, demonstrates its global receptive field in representing the content. However, the way exploiting the domain-invariant property for selective state space is rarely explored. In this paper, we propose a novel Flow Factorized State Space model, dubbed as DG-Famba, for visual domain generalization. To maintain domain consistency, we innovatively map the style-augmented and the original state embeddings by flow factorization. In this latent flow space, each state embedding from a certain style is specified by a latent probability path. By aligning these probability paths in the latent space, the state embeddings are able to represent the same content distribution regardless of the style differences. Extensive experiments conducted on various visual domain generalization settings show its state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。