arXiv:2411.16327cs.CV2024-11被引 2

用HDR图像和文字描述生成更真实的红外图像,解决细节丢失问题。

CapHDR2IR: Caption-Driven Transfer from Visible Light to Infrared Domain

论文配图:CapHDR2IR: Caption-Driven Transfer from Visible Light to Infrared Domain
图 1 · 摘自论文原文
  • 结合HDR图像与视觉语言模型,提升跨域生成能力
  • 在HDRT数据集上达到当前最佳性能,减少伪热交叉伪影
  • 适合需要高保真红外合成的遥感、安防场景

红外成像在极端光照条件下具有独特优势,但高分辨率红外传感器硬件成本高,限制了广泛应用。现有方法利用可见光合成红外图像,但在极端光照下因动态范围有限导致细节失真,并因缺乏场景上下文理解产生伪热交叉伪影。当多个温度相近物体在训练数据中无法区分时,该问题更为严重。为此,本文提出CapHDR2IR框架,采用高动态范围(HDR)图像作为输入,结合视觉语言模型中的密集文字描述分支,增强对场景语义的理解。HDR图像能覆盖更广的亮度变化,确保不同光照条件下的可靠生成;文字描述则提升了输出图像的语义一致性与可分辨性。在HDRT数据集上的大量实验表明,所提方法在可见光到红外图像转换任务中优于现有通用迁移方法及专用于该任务的方法,实现了最先进的性能。

原文摘要 · Abstract (English)

Infrared (IR) imaging offers advantages in several fields due to its unique ability of capturing content in extreme light conditions. However, the demanding hardware requirements of high-resolution IR sensors limit its widespread application. As an alternative, visible light can be used to synthesize IR images but this causes a loss of fidelity in image details and introduces inconsistencies due to lack of contextual awareness of the scene. This stems from a combination of using visible light with a standard dynamic range, especially under extreme lighting, and a lack of contextual awareness can result in pseudo-thermal-crossover artifacts. This occurs when multiple objects with similar temperatures appear indistinguishable in the training data, further exacerbating the loss of fidelity. To solve this challenge, this paper proposes CapHDR2IR, a novel framework incorporating vision-language models using high dynamic range (HDR) images as inputs to generate IR images. HDR images capture a wider range of luminance variations, ensuring reliable IR image generation in different light conditions. Additionally, a dense caption branch integrates semantic understanding, resulting in more meaningful and discernible IR outputs. Extensive experiments on the HDRT dataset show that the proposed CapHDR2IR achieves state-of-the-art performance compared with existing general domain transfer methods and those tailored for visible-to-infrared image translation.

红外生成HDR图像视觉语言模型跨域合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。