用视觉语言理解提升可见光转红外图像质量,解决语义对齐与数据稀缺难题。
DiffV2IR: Visible-to-Infrared Diffusion Model via Vision-Language Understanding
- 引入渐进式学习与视觉语言理解模块,增强跨模态语义对齐。
- 在50万张红外图像数据集上实现更真实、细节丰富的图像转换。
- 适合图像翻译、跨模态生成及红外视觉研究者使用。
可见光转红外图像(V2IR)任务因三大挑战而困难:1)实现语义感知的翻译;2)管理红外成像中多样的波长谱;3)红外数据集稀缺。现有主流方法常将V2IR视为常规图像到图像合成问题,忽略上述特性。为此,我们提出DiffV2IR框架,包含两个核心组件:渐进式学习模块(PLM)和视觉语言理解模块(VLUM)。PLM采用自适应扩散模型架构,通过多阶段知识学习实现从全波段到目标波段的红外转换。为提升翻译效果,VLUM融合统一的视觉-语言理解能力。此外,我们构建了大规模红外数据集IR-500K,包含50万张在多种场景、物体和环境条件下采集的红外图像。结合PLM、VLUM与海量数据,DiffV2IR显著提升V2IR性能。实验验证其在生成高质量图像方面的优越性,证明其有效性与广泛应用潜力。代码、数据集及模型将开源。
原文摘要 · Abstract (English)
The task of translating visible-to-infrared images (V2IR) is inherently challenging due to three main obstacles: 1) achieving semantic-aware translation, 2) managing the diverse wavelength spectrum in infrared imagery, and 3) the scarcity of comprehensive infrared datasets. Current leading methods tend to treat V2IR as a conventional image-to-image synthesis challenge, often overlooking these specific issues. To address this, we introduce DiffV2IR, a novel framework for image translation comprising two key elements: a Progressive Learning Module (PLM) and a Vision-Language Understanding Module (VLUM). PLM features an adaptive diffusion model architecture that leverages multi-stage knowledge learning to infrared transition from full-range to target wavelength. To improve V2IR translation, VLUM incorporates unified Vision-Language Understanding. We also collected a large infrared dataset, IR-500K, which includes 500,000 infrared images compiled by various scenes and objects under various environmental conditions. Through the combination of PLM, VLUM, and the extensive IR-500K dataset, DiffV2IR markedly improves the performance of V2IR. Experiments validate DiffV2IR's excellence in producing high-quality translations, establishing its efficacy and broad applicability. The code, dataset, and DiffV2IR model will be available at https://github.com/LidongWang-26/DiffV2IR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。