用视觉大模型实现高质量从可见光到近红外图像转换
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
- 基于视觉大模型的编码器-解码器结构,结合跨注意力机制融合特征
- 在RANUS数据集上FID得分提升34.81%,生成图像更真实细节更丰富
- 可有效扩充近红外数据集,适合需增强近红外视觉任务的研究者
本文提出Pix2Next,一种新颖的图像到图像翻译框架,用于从RGB输入生成高质量近红外(NIR)图像。该方法在编码器-解码器架构中引入先进的视觉基础模型(VFM),通过跨注意力机制增强特征整合能力,捕捉全局细节并保留关键光谱特性,将RGB到NIR的转换视为超越简单域迁移的问题。多尺度PatchGAN判别器在不同细节层级上确保图像真实性,而精心设计的损失函数则同时兼顾全局语义理解与局部特征保持。我们在RANUS数据集上进行了实验,结果表明Pix2Next在定量指标和视觉质量上均优于现有方法,FID得分提升34.81%。此外,我们展示了其实际应用价值:利用生成的NIR数据增强有限的真实NIR数据集后,下游目标检测任务性能得到提升。该方法无需额外采集或标注即可扩展NIR数据规模,有望加速基于近红外的计算机视觉应用发展。
原文摘要 · Abstract (English)
This paper proposes Pix2Next, a novel image-to-image translation framework designed to address the challenge of generating high-quality Near-Infrared (NIR) images from RGB inputs. Our approach leverages a state-of-the-art Vision Foundation Model (VFM) within an encoder-decoder architecture, incorporating cross-attention mechanisms to enhance feature integration. This design captures detailed global representations and preserves essential spectral characteristics, treating RGB-to-NIR translation as more than a simple domain transfer problem. A multi-scale PatchGAN discriminator ensures realistic image generation at various detail levels, while carefully designed loss functions couple global context understanding with local feature preservation. We performed experiments on the RANUS dataset to demonstrate Pix2Next's advantages in quantitative metrics and visual quality, improving the FID score by 34.81% compared to existing methods. Furthermore, we demonstrate the practical utility of Pix2Next by showing improved performance on a downstream object detection task using generated NIR data to augment limited real NIR datasets. The proposed approach enables the scaling up of NIR datasets without additional data acquisition or annotation efforts, potentially accelerating advancements in NIR-based computer vision applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。