通过多帧交叉注意力与扫描状态空间解码,提升突发图像超分辨率质量
Burst Image Super-Resolution via Multi-Cross Attention Encoding and Multi-Scan State-Space Decoding
- 设计重叠跨窗与跨帧注意力,精准提取多帧子像素信息
- 在真实与合成数据集上均超越现有方法,分辨率测试图验证效果
- 适合需要高精度图像重建的摄影、遥感等场景
多图像超分辨率(MISR)通过聚合多个空间偏移帧的亚像素信息,可实现比单图像超分辨率(SISR)更高的图像质量。其中,突发图像超分辨率(BurstSR)因应用广泛而受到广泛关注。近年来,由于在捕捉局部与全局上下文方面表现更优,Transformer逐渐取代卷积神经网络(CNN)成为超分辨率任务的主流架构。然而,大多数现有方法仍依赖固定且狭窄的注意力窗口,限制了对局部范围之外特征的感知能力,从而影响对齐与特征聚合,这两者对高质量超分辨率至关重要。为此,我们提出一种新型特征提取器,引入两种新设计的注意力机制:重叠跨窗注意力和跨帧注意力,以更精确高效地跨多帧提取亚像素信息。此外,我们还引入带有跨帧注意力机制的多扫描状态空间模块,增强特征聚合能力。在合成与真实世界基准上的大量实验表明,该方法性能优越。对ISO 12233分辨率测试图的额外评估进一步证实其显著提升的超分辨率表现。
原文摘要 · Abstract (English)
Multi-image super-resolution (MISR) can achieve higher image quality than single-image super-resolution (SISR) by aggregating sub-pixel information from multiple spatially shifted frames. Among MISR tasks, burst super-resolution (BurstSR) has gained significant attention due to its wide range of applications. Recent methods have increasingly adopted Transformers over convolutional neural networks (CNNs) in super-resolution tasks, due to their superior ability to capture both local and global context. However, most existing approaches still rely on fixed and narrow attention windows that restrict the perception of features beyond the local field. This limitation hampers alignment and feature aggregation, both of which are crucial for high-quality super-resolution. To address these limitations, we propose a novel feature extractor that incorporates two newly designed attention mechanisms: overlapping cross-window attention and cross-frame attention, enabling more precise and efficient extraction of sub-pixel information across multiple frames. Furthermore, we introduce a Multi-scan State-Space Module with the cross-frame attention mechanism to enhance feature aggregation. Extensive experiments on both synthetic and real-world benchmarks demonstrate the superiority of our approach. Additional evaluations on ISO 12233 resolution test charts further confirm its enhanced super-resolution performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。