arXiv:2601.04682cs.CV2026-01AAAI被引 3

针对红外视频湍流模糊,提出热感知扩散模型联合恢复细节与湍流失真。

HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-Resolution

论文配图:HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-Resolution
图 1 · 摘自论文原文
  • 引入热感知形变先验,通过相位引导流估计增强扩散过程的物理一致性。
  • 在自建的FLIR-IVSR数据集上,相比基线提升2.1~3.8dB PSNR,有效抑制非均匀畸变。
  • 适合关注红外视觉、湍流去模糊与扩散模型应用的研究者。

红外视频在恶劣环境视觉任务中备受关注,但常受大气湍流和压缩退化影响。现有视频超分辨率(VSR)方法或忽略红外与可见光图像的模态差异,或无法恢复湍流引起的畸变。直接串联湍流抑制(TM)与VSR方法会导致退化建模解耦,引发误差传播。本文提出HATIR,一种热感知扩散模型,用于湍流红外视频超分辨率,通过在扩散采样路径中注入热感知形变先验,联合建模湍流退化与结构细节丢失的逆过程。具体地,构建基于物理原理的相位引导流估计器,利用热活跃区域时间上一致的相位响应,实现可靠的湍流感知流引导反向扩散。为保障非均匀畸变下的结构恢复保真度,提出湍流感知解码器,通过湍流门控与结构感知注意力,选择性抑制不稳定的时序信号并增强边缘特征聚合。构建了首个湍流红外视频超分辨率数据集FLIR-IVSR,包含640个不同场景的配对低/高分辨率序列,由FLIR T1050sc相机(1024×768)采集,涵盖多样的相机与物体运动条件,推动该领域研究发展。

原文摘要 · Abstract (English)

Infrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fail to restore turbulence-induced distortions. Directly cascading turbulence mitigation (TM) algorithms with VSR methods leads to error propagation and accumulation due to the decoupled modeling of degradation between turbulence and resolution. We introduce HATIR, a Heat-Aware Diffusion for Turbulent InfraRed Video Super-Resolution, which injects heat-aware deformation priors into the diffusion sampling path to jointly model the inverse process of turbulent degradation and structural detail loss. Specifically, HATIR constructs a Phasor-Guided Flow Estimator, rooted in the physical principle that thermally active regions exhibit consistent phasor responses over time, enabling reliable turbulence-aware flow to guide the reverse diffusion process. To ensure the fidelity of structural recovery under nonuniform distortions, a Turbulence-Aware Decoder is proposed to selectively suppress unstable temporal cues and enhance edge-aware feature aggregation via turbulence gating and structure-aware attention. We built FLIR-IVSR, the first dataset for turbulent infrared VSR, comprising paired LR-HR sequences from a FLIR T1050sc camera (1024 X 768) spanning 640 diverse scenes with varying camera and object motion conditions. This encourages future research in infrared VSR. Project page: https://github.com/JZ0606/HATIR

红外视频超分辨率扩散模型湍流去模糊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。