让图像翻译模型学会识别配对质量,自动过滤错误对应。
A${}^2$BM: Alignment-Aware Bridge Matching for Image-to-Image Translation

- 引入对齐评分机制,在训练中区分真实对应与错位干扰。
- 在跨传感器超分和无监督域适应任务上,性能超越主流基线。
- 适合处理采集条件差异大的真实场景图像翻译问题。
成对图像到图像翻译支撑图像编辑、传感器转换和域自适应等多种视觉任务。桥接匹配与流匹配作为新兴框架,将扩散模型扩展至任意源-目标分布。但其标准形式假设训练对完全对齐,认为所有源-目标对应均可靠。现实中,因异步捕获、光照差异或配准误差,常出现弱对齐数据。本文提出对齐感知桥接匹配(A²BM),在训练中利用图像对的对齐评分,使模型学会剥离真实语义对应与错位伪影。推理时,以对齐评分为控制变量调节翻译保真度,高分对应输出更精确。我们在合成实验和真实挑战任务(如跨传感器超分辨率、像素空间无监督域适应)上验证,A²BM 在所有设置下均显著优于强基线(包括GAN、扩散模型及Schrödinger桥方法),确立对齐条件化为弱对齐数据下图像翻译的合理解决方案。
原文摘要 · Abstract (English)
Paired image-to-image translation underpins a wide range of computer vision tasks, including image editing, sensor translation, and domain adaptation. Bridge matching and flow matching have recently emerged as powerful frameworks, extending diffusion models to arbitrary source and target distributions. However, their standard formulations assume perfectly aligned training pairs, treating all source-target correspondences as equally reliable. In practice, real-world applications often involve weakly aligned pairs due to changes of acquisition conditions, including e.g. asynchronous captures, different illuminations, or misregistration. In this work, we introduce Alignment-Aware Bridge Matching (A${}^2$BM), a bridge matching method that leverages image pairs alignment during training. By incorporating alignment scores, the model learns to disentangle true semantic correspondences from misalignment artifacts. At inference time, we use the alignment score as a control variable over translation fidelity, with strongly aligned outputs obtained when prompting the model with the highest alignment score. We validate A${}^2$BM on both controlled synthetic experiments and on challenging real-world tasks, including cross-sensor super-resolution and pixel-space unsupervised domain adaptation. In all settings, A${}^2$BM consistently improves translation fidelity over strong GAN-, diffusion-, and Schr{ö}dinger bridge-based baselines, establishing alignment conditioning as a principled solution for image translation models with weakly aligned data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。