全面评估生成式图像修复的进展与瓶颈,揭示从细节缺失到过生成的新挑战。
How far have we gone in Generative Image Restoration? A study on its capability, limitations and evaluation practices
- 构建多维评估框架,量化分析模型在细节、清晰度、语义正确性等方面表现。
- 发现当前模型普遍面临过生成问题,而非早期的细节不足缺陷。
- 提出更贴近人类感知的图像质量评价模型,适用于未来研究参考。
生成式图像修复(GIR)在感知真实感方面取得了显著进展,但其实际能力相较于传统方法究竟提升了多少?为此,我们提出了一个大规模研究,基于全新的多维度评估流程,对模型在细节、锐度、语义正确性和整体质量上的表现进行系统评估。研究涵盖扩散模型、GAN模型、基于PSNR优化的模型以及通用生成模型等多种架构,揭示了显著的性能差异。此外,我们的分析发现失败模式出现关键演变,标志着以感知为导向的低层视觉领域发生范式转变:核心挑战已从过去的细节缺失(欠生成)演变为当前的细节质量与语义控制(过生成)。我们还利用该基准训练了一个新型图像质量评价(IQA)模型,其评价结果更符合人类感知判断。本工作为现代生成式图像修复模型提供了系统性研究,重新定义了对其真实水平的理解,并为未来发展方向提供了重要指引。
原文摘要 · Abstract (English)
Generative Image Restoration (GIR) has achieved impressive perceptual realism, but how far have its practical capabilities truly advanced compared with previous methods? To answer this, we present a large-scale study grounded in a new multi-dimensional evaluation pipeline that assesses models on detail, sharpness, semantic correctness, and overall quality. Our analysis covers diverse architectures, including diffusion-based, GAN-based, PSNR-oriented, and general-purpose generation models, revealing critical performance disparities. Furthermore, our analysis uncovers a key evolution in failure modes that signifies a paradigm shift for the perception-oriented low-level vision field. The central challenge is evolving from the previous problem of detail scarcity (under-generation) to the new frontier of detail quality and semantic control (preventing over-generation). We also leverage our benchmark to train a new IQA model that better aligns with human perceptual judgments. Ultimately, this work provides a systematic study of modern generative image restoration models, offering crucial insights that redefine our understanding of their true state and chart a course for future development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。