用统一模型实现红外与可见光图像双向生成,提升跨模态翻译质量。
CM-Diff: A Single Generative Network for Bidirectional Cross-Modality Translation Diffusion Model Between Infrared and Visible Images
- 通过双向扩散训练学习双模态数据分布,引入方向标签引导生成。
- 在多个数据集上优于现有方法,生成图像更贴近目标模态真实分布。
- 适合需要双模态数据增强的红外/可见光视觉任务研究者使用。
红外与可见光图像之间的图像翻译是缓解模态信息缺失的关键方法,同时有助于增强特定模态的数据集。然而,现有方法或仅支持单向翻译,或依赖循环一致性实现双向翻译,可能导致性能不佳。本文提出双向跨模态翻译扩散模型(CM-Diff),可同时建模红外与可见光模态的数据分布。通过引入翻译方向标签进行训练引导,并结合跨模态特征控制,将模态映射关系视为学习数据分布与理解模态差异的过程,提出新颖的双向扩散训练(BDT)。此外,设计统计约束推理(SCI)机制,确保生成图像紧密遵循目标模态的数据分布。实验结果表明,CM-Diff在多个基准上超越当前最优方法,展现出生成双模态数据集的巨大潜力。
原文摘要 · Abstract (English)
Image translation is one of the crucial approaches for mitigating information deficiencies in the infrared and visible modalities, while also facilitating the enhancement of modality-specific datasets. However, existing methods for infrared and visible image translation either achieve unidirectional modality translation or rely on cycle consistency for bidirectional modality translation, which may result in suboptimal performance. In this work, we present the bidirectional cross-modality translation diffusion model (CM-Diff) for simultaneously modeling data distributions in both the infrared and visible modalities. We address this challenge by combining translation direction labels for guidance during training with cross-modality feature control. Specifically, we view the establishment of the mapping relationship between the two modalities as the process of learning data distributions and understanding modality differences, achieved through a novel Bidirectional Diffusion Training (BDT). Additionally, we propose a Statistical Constraint Inference (SCI) to ensure the generated image closely adheres to the data distribution of the target modality. Experimental results demonstrate the superiority of our CM-Diff over state-of-the-art methods, highlighting its potential for generating dual-modality datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。