arXiv:2503.01187cs.CV2025-03CVPR被引 35

用梯度引导扩散模型提升红外图像超分辨率,兼顾视觉质量与下游任务表现。

DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution

论文配图:DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution
图 1 · 摘自论文原文
  • 通过注入视觉与感知先验梯度,实现逆向生成过程的任务导向优化。
  • 在IRSTD-1k数据集上,峰值信噪比达34.82,检测与分割任务性能领先。
  • 适合需要高保真红外图像的自动驾驶与机器人感知系统使用。

红外成像在自动驾驶和机器人操作中至关重要,因其在复杂环境下的稳定表现。然而,红外相机普遍存在空间分辨率低、退化复杂等问题,严重影响成像质量和后续视觉任务。为此,红外图像超分辨率(IISR)被提出以应对挑战。尽管扩散模型近年取得显著进展,现有方法或忽略红外成像的模态特性,或忽视机器感知需求。为此,我们提出DifIISR,一种面向视觉质量与感知性能优化的红外图像超分辨率扩散模型。通过在反向生成过程中注入基于视觉与感知先验的梯度,实现任务导向引导。具体而言,引入红外热谱分布调控机制,保留视觉保真度,使重建图像在频域成分上与高分辨率图像高度匹配。随后,融合多种视觉基础模型作为感知引导,注入可泛化的感知特征,有益于检测与分割任务。实验表明,该方法不仅在视觉质量上表现优异,且在下游任务中达到最新水平。代码已开源:https://github.com/zirui0625/DifIISR。

原文摘要 · Abstract (English)

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challenge imaging quality and subsequent visual tasks. Hence, infrared image super-resolution (IISR) has been developed to address this challenge. While recent developments in diffusion models have greatly advanced this field, current methods to solve it either ignore the unique modal characteristics of infrared imaging or overlook the machine perception requirements. To bridge these gaps, we propose DifIISR, an infrared image super-resolution diffusion model optimized for visual quality and perceptual performance. Our approach achieves task-based guidance for diffusion by injecting gradients derived from visual and perceptual priors into the noise during the reverse process. Specifically, we introduce an infrared thermal spectrum distribution regulation to preserve visual fidelity, ensuring that the reconstructed infrared images closely align with high-resolution images by matching their frequency components. Subsequently, we incorporate various visual foundational models as the perceptual guidance for downstream visual tasks, infusing generalizable perceptual features beneficial for detection and segmentation. As a result, our approach gains superior visual results while attaining State-Of-The-Art downstream task performance. Code is available at https://github.com/zirui0625/DifIISR

红外图像超分辨率扩散模型感知引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。