用离散代码本学习提升模型跨域泛化能力
Domain Generalization via Discrete Codebook Learning
- 将特征图量化为离散码字,聚焦语义信息而非像素细节
- 在多个基准上优于现有方法,显著缩小领域间分布差距
- 适合需要强泛化能力的跨域场景,如医疗图像分析
领域泛化(DG)旨在应对不同环境间的分布偏移,提升模型泛化能力。现有方法局限于连续特征表示的学习,通常在像素级进行训练。然而,这种范式在处理大规模连续特征空间时难以缓解分布差异,易受像素级伪相关或噪声影响。本文首次从理论上证明,离散化过程可减小连续表示学习中的领域差距。基于此,提出一种新的DG学习范式——离散领域泛化(DDG)。DDG利用代码本将特征图量化为离散码字,在共享的离散表示空间中对齐语义等价信息,优先保留语义层级信息,弱化像素级细节。通过语义级学习,减少潜在特征数量,优化表示空间利用,降低连续特征空间过大带来的风险。在广泛使用的多个DG基准上,实验表明DDG显著优于现有先进方法,验证其有效缩小分布差距、增强模型泛化能力的潜力。
原文摘要 · Abstract (English)
Domain generalization (DG) strives to address distribution shifts across diverse environments to enhance model's generalizability. Current DG approaches are confined to acquiring robust representations with continuous features, specifically training at the pixel level. However, this DG paradigm may struggle to mitigate distribution gaps in dealing with a large space of continuous features, rendering it susceptible to pixel details that exhibit spurious correlations or noise. In this paper, we first theoretically demonstrate that the domain gaps in continuous representation learning can be reduced by the discretization process. Based on this inspiring finding, we introduce a novel learning paradigm for DG, termed Discrete Domain Generalization (DDG). DDG proposes to use a codebook to quantize the feature map into discrete codewords, aligning semantic-equivalent information in a shared discrete representation space that prioritizes semantic-level information over pixel-level intricacies. By learning at the semantic level, DDG diminishes the number of latent features, optimizing the utilization of the representation space and alleviating the risks associated with the wide-ranging space of continuous features. Extensive experiments across widely employed benchmarks in DG demonstrate DDG's superior performance compared to state-of-the-art approaches, underscoring its potential to reduce the distribution gaps and enhance the model's generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。