针对红外图像超分,提出基于区域先验的注意力机制,提升固定视角场景下的重建效果
RPT-SR: Regional Prior attention Transformer for infrared image Super-Resolution
- 引入双令牌结构,用全局场景先验与局部内容结合建模
- 在LWIR/SWIR多数据集上实现最新性能,峰值达42.3dB
- 适合安防、自动驾驶等固定视角红外成像任务
通用超分辨率模型,尤其是视觉变压器,在监控和自动驾驶等固定或近似静态视角的红外成像场景中表现不佳,因其无法利用场景中固有的强而持久的空间先验信息,导致冗余学习和性能不足。为此,我们提出面向红外图像超分辨率的区域先验注意力变换器(RPT-SR),通过显式将场景布局信息编码进注意力机制来解决该问题。核心贡献是双令牌框架:融合可学习的区域先验令牌(作为场景全局结构的持久记忆)与捕捉当前输入帧特定内容的局部令牌。通过将这些令牌引入注意力机制,使先验信息动态调制局部重建过程。大量实验验证了方法的有效性。与以往研究多聚焦单一红外波段不同,RPT-SR在涵盖长波(LWIR)和短波(SWIR)谱段的多个数据集上均取得新的最先进性能。
原文摘要 · Abstract (English)
General-purpose super-resolution models, particularly Vision Transformers, have achieved remarkable success but exhibit fundamental inefficiencies in common infrared imaging scenarios like surveillance and autonomous driving, which operate from fixed or nearly-static viewpoints. These models fail to exploit the strong, persistent spatial priors inherent in such scenes, leading to redundant learning and suboptimal performance. To address this, we propose the Regional Prior attention Transformer for infrared image Super-Resolution (RPT-SR), a novel architecture that explicitly encodes scene layout information into the attention mechanism. Our core contribution is a dual-token framework that fuses (1) learnable, regional prior tokens, which act as a persistent memory for the scene's global structure, with (2) local tokens that capture the frame-specific content of the current input. By utilizing these tokens into an attention, our model allows the priors to dynamically modulate the local reconstruction process. Extensive experiments validate our approach. While most prior works focus on a single infrared band, we demonstrate the broad applicability and versatility of RPT-SR by establishing new state-of-the-art performance across diverse datasets covering both Long-Wave (LWIR) and Short-Wave (SWIR) spectra
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。