arXiv:2503.17825cs.CV2025-03被引 3

用分形结构高效提升多种图像修复效果,无需复杂自注意力机制。

Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration

  • 通过分形设计逐级扩展局部信息,融合细节与全局上下文。
  • 在7类修复任务中达顶尖性能,2倍超分提升0.21 dB PSNR。
  • 适合追求高效高保真图像修复的开发者和研究者。

尽管视觉变换器在各类图像修复(IR)任务中取得显著进展,但在多类型退化和不同分辨率下实现高效扩展仍具挑战。本文提出Fractal-IR,一种基于分形的设计,通过反复将局部信息拓展至更大区域,逐步优化退化图像。该架构在早期自然捕捉局部细节,深层阶段无缝过渡到全局上下文,避免了计算量大的长距离自注意力机制。此外,我们分析了视觉变换器在图像修复中扩展的难题,提出一套系统性策略以有效指导模型扩展。大量实验表明,Fractal-IR在七种常见图像修复任务中均达到当前最优表现,包括超分辨率、去噪、JPEG伪影去除、恶劣天气下修复、运动模糊去模糊、散焦去模糊及去马赛克。在Manga109数据集上,2×超分辨率任务中,其PSNR比之前方法提升0.21 dB;在Urban100数据集上,对于σ=50的灰度图像去噪,性能领先0.2 dB。

原文摘要 · Abstract (English)

While vision transformers achieve significant breakthroughs in various image restoration (IR) tasks, it is still challenging to efficiently scale them across multiple types of degradations and resolutions. In this paper, we propose Fractal-IR, a fractal-based design that progressively refines degraded images by repeatedly expanding local information into broader regions. This fractal architecture naturally captures local details at early stages and seamlessly transitions toward global context in deeper fractal stages, removing the need for computationally heavy long-range self-attention mechanisms. Moveover, we observe the challenge in scaling up vision transformers for IR tasks. Through a series of analyses, we identify a holistic set of strategies to effectively guide model scaling. Extensive experimental results show that Fractal-IR achieves state-of-the-art performance in seven common image restoration tasks, including super-resolution, denoising, JPEG artifact removal, IR in adverse weather conditions, motion deblurring, defocus deblurring, and demosaicking. For $2\times$ SR on Manga109, Fractal-IR achieves a 0.21 dB PSNR gain. For grayscale image denoising on Urban100, Fractal-IR surpasses the previous method by 0.2 dB for $σ=50$.

图像修复分形结构视觉变换器超分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。