提出分层信息流机制,让模型高效通用地处理多种图像修复任务。
Hierarchical Information Flow for Generalized Efficient Image Restoration
- 自下而上构建分层信息树,逐级融合局部细节与全局语义。
- 在7个常见修复任务上达顶尖性能,且计算效率更高。
- 适合需要高效、可扩展的通用图像修复系统的研究者使用。
尽管视觉变换器在多种图像修复(IR)任务中表现优异,但如何高效地泛化并扩展模型以适应多任务仍具挑战。为平衡效率与模型容量,本文提出一种面向图像修复的分层信息流机制——Hi-IR,其通过自底向上的方式逐步传播像素间信息。Hi-IR 构建了包含三个层级的分层信息树,分别表示退化图像的不同层次信息:高层聚焦于整体对象与概念,低层关注局部细节。该架构摒弃长距离自注意力,提升计算效率与内存利用率,有利于模型大规模扩展。基于此,我们探索模型缩放策略,增强方法能力,适用于大规模训练场景。大量实验表明,Hi-IR 在七个常见的图像修复任务中均达到领先性能,验证了其有效性与泛化能力。
原文摘要 · Abstract (English)
While vision transformers show promise in numerous image restoration (IR) tasks, the challenge remains in efficiently generalizing and scaling up a model for multiple IR tasks. To strike a balance between efficiency and model capacity for a generalized transformer-based IR method, we propose a hierarchical information flow mechanism for image restoration, dubbed Hi-IR, which progressively propagates information among pixels in a bottom-up manner. Hi-IR constructs a hierarchical information tree representing the degraded image across three levels. Each level encapsulates different types of information, with higher levels encompassing broader objects and concepts and lower levels focusing on local details. Moreover, the hierarchical tree architecture removes long-range self-attention, improves the computational efficiency and memory utilization, thus preparing it for effective model scaling. Based on that, we explore model scaling to improve our method's capabilities, which is expected to positively impact IR in large-scale training settings. Extensive experimental results show that Hi-IR achieves state-of-the-art performance in seven common image restoration tasks, affirming its effectiveness and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。