用GPT生成图像修复,虽视觉好看但结构不稳,却可提升现有模型效果。
A Preliminary Study on GPT-Image Generation Model for Image Restoration
- 用GPT生成图像作为先验,增强修复模型表现
- 生成结果感知好但像素级结构常失真
- 适合想融合生成模型的图像修复研究者
近期OpenAI GPT系列多模态生成模型在生成视觉吸引人图像方面表现出色。本文首次系统评估其在图像修复任务中的潜力。实验表明,尽管GPT-Image生成结果在感知上令人愉悦,但与真实参考图相比,往往缺乏像素级结构保真度,常见偏差包括图像几何、物体位置或数量变化,甚至视角改变。此外,我们进一步证明GPT-Image输出可作为强大视觉先验,显著提升现有修复网络性能。以去雾、去雨、低光增强为例,融入GPT生成先验后修复质量明显改善。本研究为将GPT类生成先验引入修复流程提供实用洞见与基线框架,也揭示了生成模型与修复任务结合的新机遇。为支持后续研究,我们将公开GPT修复结果。
原文摘要 · Abstract (English)
Recent advances in OpenAI's GPT-series multimodal generation models have shown remarkable capabilities in producing visually compelling images. In this work, we investigate its potential impact on the image restoration community. We provide, to the best of our knowledge, the first systematic benchmark across diverse restoration scenarios. Our evaluation shows that, while the restoration results generated by GPT-Image models are often perceptually pleasant, they tend to lack pixel-level structural fidelity compared with ground-truth references. Typical deviations include changes in image geometry, object positions or counts, and even modifications in perspective. Beyond empirical observations, we further demonstrate that outputs from GPT-Image models can act as strong visual priors, offering notable performance improvements for existing restoration networks. Using dehazing, deraining, and low-light enhancement as representative case studies, we show that integrating GPT-generated priors significantly boosts restoration quality. This study not only provides practical insights and a baseline framework for incorporating GPT-based generative priors into restoration pipelines, but also highlights new opportunities for bridging image generation models and restoration tasks. To support future research, we will release GPT-restored results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。