用自监督语义桥实现无需配对的高保真图像翻译
Unpaired Image-to-Image Translation via a Self-Supervised Semantic Bridge
- 通过自监督视觉编码器构建几何不变的共享潜在空间
- 在医学图像合成任务中显著提升域内与域外泛化性能
- 支持高质量文本引导编辑,无需跨域标注
对抗性扩散与扩散反演方法虽推动了无配对图像翻译发展,但各有局限:前者训练需目标域对抗损失,影响泛化能力;后者常因噪声潜在表示重建不完美导致生成质量低下。本文提出自监督语义桥(SSB)框架,将外部语义先验融入扩散桥模型,实现无需跨域监督的空间保真翻译。核心思想是利用自监督视觉编码器学习对外观变化不变但保留几何结构的表征,构建共享潜在空间以指导扩散过程。大量实验表明,SSB在医学图像合成任务中优于现有强基线方法,无论域内还是域外设置均表现优异,并可轻松扩展至高质量文本引导编辑。
原文摘要 · Abstract (English)
Adversarial diffusion and diffusion-inversion methods have advanced unpaired image-to-image translation, but each faces key limitations. Adversarial approaches require target-domain adversarial loss during training, which can limit generalization to unseen data, while diffusion-inversion methods often produce low-fidelity translations due to imperfect inversion into noise-latent representations. In this work, we propose the Self-Supervised Semantic Bridge (SSB), a versatile framework that integrates external semantic priors into diffusion bridge models to enable spatially faithful translation without cross-domain supervision. Our key idea is to leverage self-supervised visual encoders to learn representations that are invariant to appearance changes but capture geometric structure, forming a shared latent space that conditions the diffusion bridges. Extensive experiments show that SSB outperforms strong prior methods for challenging medical image synthesis in both in-domain and out-of-domain settings, and extends easily to high-quality text-guided editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。