通过连续巴氏中心空间学习统一表征,提升图像修复模型对未知退化的泛化能力。
BaryIR: Learning Multi-Source Unified Representation in Continuous Barycenter Space for Generalizable All-in-One Image Restoration
- 构建连续巴氏中心空间,融合多源退化特征的统一表征
- 在真实数据和未见退化上表现更优,优于现有顶尖方法
- 适合需要跨退化场景通用修复的应用场景
尽管所有类型图像修复(AIR)在同时处理多种退化方面取得了显著进展,但现有方法对分布外退化和图像仍敏感,限制了实际应用。本文提出多源表示学习框架BaryIR,将多源退化图像的隐空间分解为连续巴氏中心空间以实现统一特征编码,以及源特定子空间以进行特定语义编码。具体地,引入多源潜在最优传输巴氏中心问题,学习连续巴氏中心映射以将潜在表示迁移到巴氏中心空间。传输代价设计使源特定子空间的表示相互对比,同时与巴氏中心空间保持正交性。这使得BaryIR能在巴氏中心空间中学习紧凑的退化无关信息,并在源特定子空间中保留退化相关语义,捕捉多源数据流形的内在几何结构,实现可泛化的所有类型图像修复。大量实验表明,BaryIR在性能上媲美当前最优所有类型方法,尤其在真实世界数据和未见退化上表现出更强的泛化能力。代码将在https://github.com/xl-tang3/BaryIR公开。
原文摘要 · Abstract (English)
Despite remarkable advances made in all-in-one image restoration (AIR) for handling different types of degradations simultaneously, existing methods remain vulnerable to out-of-distribution degradations and images, limiting their real-world applicability. In this paper, we propose a multi-source representation learning framework BaryIR, which decomposes the latent space of multi-source degraded images into a continuous barycenter space for unified feature encoding and source-specific subspaces for specific semantic encoding. Specifically, we seek the multi-source unified representation by introducing a multi-source latent optimal transport barycenter problem, in which a continuous barycenter map is learned to transport the latent representations to the barycenter space. The transport cost is designed such that the representations from source-specific subspaces are contrasted with each other while maintaining orthogonality to those from the barycenter space. This enables BaryIR to learn compact representations with unified degradation-agnostic information from the barycenter space, as well as degradation-specific semantics from source-specific subspaces, capturing the inherent geometry of multi-source data manifold for generalizable AIR. Extensive experiments demonstrate that BaryIR achieves competitive performance compared to state-of-the-art all-in-one methods. Particularly, BaryIR exhibits superior generalization ability to real-world data and unseen degradations. The code will be publicly available at https://github.com/xl-tang3/BaryIR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。